Broadcast Prism One story, split by channel. See the full spectrum.
Home Stories Broadcasts Channels About

Legal

Methodology

What every number on this site means, how it's calculated, and when to trust it.

Last updated: 17 August 2026

Broadcast Prism turns TV news transcripts into a map of attention — what each channel covered, for how long, and where agendas lined up or diverged. This page explains every metric on the site: what it measures, how it's calculated, where it can mislead, and the minimum sample size before it's worth trusting. These are measures of attention, not verdicts on truth.

On this page

  • Story & claim clustering
  • Airtime
  • Coverage gaps & blind spots
  • Editorial fingerprints
  • Agenda alignment
  • Topic pairings & editorial network
  • Narrative synthesis
  • Story lifecycle badges
  • Attention volatility
  • Contested & charged quotes
  • Category classification
  • Broadcast stills
  • News density
  • Signature vocabulary
  • Story tone: how bleak, how hard
  • Weekly overview & superlatives
  • Web article context & source spectrum
  • Trending & new story rails
  • Lead story comparison

Story & claim clustering

What is measured. The grouping of similar reports across different TV channels into a single canonical story, and the consolidation of individual factual claims within those stories so the same claim made across channels lines up.

How it's calculated. Reports are first grouped by shared entities and semantic similarity within a single topic. A language-model deduplication pass maps provisional story hooks to active stories or creates new ones. A second, global cross-topic merge pass (default 14 days lookback) collapses stories describing the same real-world development across different topics. Finally, a cross-topic claim merge identifies and merges duplicate factual claims across the unified story.

Known failure modes. Two genuinely different stories might be merged if they share similar entities (like generic names or places), or a single story might be split if channels report it using completely different details or vocabulary. Claim merging may fail to group statements that differ slightly in nuance.

Minimum sample size. Grouping has no meaning for a single channel's isolated report; it requires coverage from at least two channels to trace alignment and overlap.

Headlines and URLs. A story's headline is rewritten as its coverage accumulates, so a story that runs for days reflects where it has got to rather than how it first broke. Its web address does not change with it: the address is fixed when the story is created, from that first framing, so a link shared on day one still works on day twenty. On a story that turned, the address can therefore read like an older version of the headline above it. Where a headline has not yet been refreshed, the original title is shown.

Airtime

What is measured. The amount of time, in seconds, a channel spent on a story or topic in a broadcast.

How it's calculated. Airtime is summed per channel from the durations of identified segments. Segment boundaries and durations are extracted from the raw transcripts based on the start and end timecodes of the speech segments associated with a story.

Known failure modes. Mis-segmented transcript boundaries can cut off sections or include unrelated content; ads and handovers may be misidentified as story content; a single topic split across non-adjacent segments may have its airtime under-counted or double-counted depending on segment alignment.

Minimum sample size. A single broadcast's airtime is noisy and susceptible to format variation (e.g. short summaries vs long packages); stable comparison of channel priorities requires a sample of several days.

Coverage gaps & blind spots

What is measured. Stories some channels covered while others omitted, flagged with HIGH, MEDIUM, or LOW severity, along with "persistent blind spots" that track these omissions over time.

How it's calculated. Omissions are detected by comparing the presence and absence of stories across the channel-by-channel matrix. Severity is scored based on the overall duration of the story and the coverage capacity of the channel. Blind spots aggregate these omissions into all-time channel-specific profiles: a category becomes a persistent blind spot when the channel accounts for at least two omissions in it and at least 35% of all omissions recorded in that category. The channels table's "Flagged omissions" column counts the omissions themselves, not the categories, and is shown beside the days the channel was on air. It is deliberately a count rather than a rate: the detection step already declines to flag an omission that a shorter bulletin explains, so dividing by airtime or by days would apply that same correction a second time. A channel with a very low count is often one whose bulletin is short enough that few omissions ever qualify, rather than one that misses little.

Known failure modes. Shorter broadcasts naturally cover fewer stories, which the algorithms try to account for but can still misidentify as gaps. A channel covering a story under a totally different frame or topic title may be incorrectly flagged as having omitted the story. Omission counts also rise with how many days a channel was on air, so they compare channels of similar output better than they compare a rolling news channel with a single nightly programme.

Minimum sample size. Individual daily gap flags are indicators of difference, not bias; identifying a systemic blind spot requires tracking coverage over a sustained period of at least two to three weeks.

Editorial fingerprints

What is measured. The unique focus areas of a channel—specifically which topics it covers significantly more or less than its peers, and which stories it alone covered.

How it's calculated. Exclusivity and emphasis are measured separately, and on different units. Exclusivity is measured on stories: a story is called exclusive to a channel when no other channel has broadcast it on any day of its run, and it received at least a minute of that channel's airtime on the day shown. Emphasis is measured on topics, and only on topics at least half the channels on air that day covered: for each of those, we take the channel's share of its own daily airtime and compare it to the average share among the channels that covered that topic, and a channel "over-indexes" where its share exceeds that average. The channels page condenses this into a distinctiveness band: we take the root-mean-square of a channel's category-share deltas from the network average (in percentage points), shrink it toward the average in proportion to how few topics we observed that channel covering, and label the result Typical below 2.5, Distinct from 2.5 to 5, and Highly distinct at 5 or above.

Known failure modes. A single massive breaking story can dominate a channel's daily airtime, throwing off fingerprints. Ratios can also be highly volatile for topics with very low absolute airtime (e.g., under 10 seconds). On days when no topic reached half the channels there is no shared baseline to compare against, so no over-index is reported at all. Exclusivity was previously measured on topic labels rather than stories, which overstated it: one story is often split across several topic names, each landing on a different channel. Measuring it on stories removes that error but not the related one — where a single subject has been split into two stories, both can be reported as exclusives. The distinctiveness band was also computed without regard to sample size, which rewarded narrowness: a single twenty-minute nightly programme spends its whole airtime on two or three topics and is scored as maximally unlike the network, both for what it over-covers and for the categories it has no room to touch. Shrinkage corrects for that but cannot manufacture confidence — a channel observed on few topics has a score that is mostly the network average, by design. The band thresholds are editorial calibrations against the live dataset, not statistical significance levels.

Minimum sample size. A single day's fingerprint can be heavily skewed by one broadcast's running order; identifying a channel's true editorial signature requires analysing several days of coverage.

Agenda alignment

What is measured. The degree of similarity between the daily news agendas of any two channels—showing which outlets align in their story selection and which diverge.

How it's calculated. We represent each channel's daily agenda as a vector of its topic airtime shares, then calculate the cosine similarity between these vectors. A score of 100% indicates identical relative emphasis across all topics, while 0% indicates no overlapping topics.

Known failure modes. Agenda alignment measures focus, not viewpoint. Two channels that spend half their broadcasts covering the same political debate will score as highly aligned, even if their reporting frames the debate in opposite ways. Low-volume days also make similarity scores highly sensitive.

Minimum sample size. Unreliable for broadcasts covering fewer than three topics, as sparse vectors artificially inflate similarity scores; requires a normal multi-story news day to be meaningful.

Topic pairings & editorial network

What is measured. For each channel, the two topics it most often runs on the same day compared with how often the rest of the roster pairs those same two — and, separately, the layout of broadcaster relationships based on shared coverage patterns.

How it's calculated. For every pair of topics on a given day we record whether a channel covered at least one of them and whether it covered both. A channel's rate for a pair is the days it ran both over the days it ran either. Ranking that rate directly would just re-sort the roster by breadth, because a channel covering fifteen topics a day runs any two of them together far more often than one covering six — so each rate is divided by that channel's own overall pairing propensity across all pairs, and then set against the same figure for the rest of the roster. The pair with the highest resulting index is the one shown. A pairing must appear on at least 4 days, from at least 8 days on which the channel ran either topic, before it qualifies. The displayed percentages are the raw rates with their supporting day counts; only the ranking is smoothed. For the broadcaster network graph, we construct undirected links between channels whose all-time cosine similarity is at least 20%.

Known failure modes. Same-day co-occurrence describes what shared a bulletin day, not why: two topics can pair because they are genuinely connected, because one caused the other, or because they simply broke in the same week. The measure is comparative, so it says a pairing is unusual for this channel against this roster — not that the channel is preoccupied with it in absolute terms. An earlier version ranked pairs by Jaccard overlap across the whole network, which mostly reported how common each topic was and promoted rare topics that happened to coincide; it also made no claim about any channel.

Minimum sample size. The 4-day and 8-day floors are the gate. Channels that do not clear them are named beneath the section rather than dropped, so their absence is visible rather than silent.

Narrative synthesis

What is measured. The neutral, cross-channel summary and channel-by-channel breakdown generated for major news stories covered by multiple broadcasters.

How it's calculated. We gather all covering channels' transcripts for a given story, clamp them to the story's time window (up to a character limit), and pass them to a large language model, prompted as an expert news analyst and editor to write a balanced summary and map individual channel angles.

Known failure modes. The model can smooth over genuine editorial disagreements into a false consensus, inherit errors or misattributions from raw transcripts, or exhibit hallucination, especially on complex or fast-moving developments. It is not manually fact-checked.

Minimum sample size. Summaries are only generated for stories covered by at least two distinct channels; single-channel reports do not qualify for cross-channel synthesis.

Story lifecycle badges

What is measured. The active status of a topic or story in the news cycle, classified into one of five states: Emerging, Peaking, Fading, Recurring, or Dormant.

How it's calculated. Status is classified deterministically by evaluating the topic's relative airtime history. It is Fading when latest airtime falls below 0.5 × its all-time max; Peaking when latest is at least 0.8 × max and greater than its previous appearance; Recurring when it has appeared on at least three distinct dates without peaking or fading; Emerging when it is a new topic or growing in airtime; and Dormant when it last appeared before the most recent complete day.

Known failure modes. A sudden one-day spike can prematurely label a topic as Peaking, even if it immediately disappears the next day. The boundaries are relative to each topic's own history, meaning low-volume topics can trigger state changes as easily as major stories. Dormancy was previously judged against the newest date in the archive, which is normally a day still being analysed — so stories that had led the previous day's coverage were labelled Dormant until the current day's bulletins came in.

Minimum sample size. A topic must appear on at least three distinct dates to qualify for Recurring status; prior to that, it will default to Emerging or Dormant.

Attention volatility

What is measured. How much a channel's daily agenda swings from one day to the next, indicating whether it maintains a stable focus or reactively shifts its priorities.

How it's calculated. We compute the cosine distance (1 minus cosine similarity) between a channel's category airtime shares on consecutive days. The average of these day-on-day distances defines its volatility score: Stable (under 25%), Moderate (25% to 45%), or Reactive (over 45%).

Known failure modes. A genuinely fast-moving, high-impact news week will cause even normally consistent channels to score as highly volatile; the score does not distinguish between chasing sensations and covering major daily developments.

Minimum sample size. Requires at least two days of coverage for a channel to compute a single day-on-day shift, and a minimum of 7 consecutive days to establish a reliable average.

Contested & charged quotes

What is measured. Direct quotations that are flagged as contested (challenged by other sources or fact-checkers) or loaded/emotionally charged, mapped directly to their surrounding transcript excerpt.

How it's calculated. A language model extracts candidate quotes from transcripts based on loaded language or claims under dispute. A verification pass matches these candidates back to the exact transcript segments to establish evidence anchors, timecodes, and still images, showing the highlighted span in its exact context. Trend graphs normalise the quote rate by dividing counts by the channel's analysed hours on that day. The channels page bands the same rate across the whole archive into an intensity chip: none at zero, low below 2 per hour, moderate below 3.5, high at 3.5 or above.

Known failure modes. Sarcasm, hypothetical statements, or paraphrases can be misidentified as loaded language; matching can fail if the transcript contains minor phonetic errors. Normalisation can also produce high rates on days with very short broadcasts if a single quote is flagged. The intensity chip previously banded raw counts rather than a rate, which measured how much of a channel we had captured rather than how it spoke, and put every channel in the top band as the archive grew. The thresholds are editorial calibrations against the observed spread, not significance levels.

Minimum sample size. Single flagged quotes are illustrative examples of framing, not statistical proof; drawing conclusions about a channel's overall quote profile requires tracking quote rates over at least a week.

Category classification

What is measured. The thematic category (such as Politics, Health, Economy, Ukraine, or Royal) a story is filed under.

How it's calculated. Assigned by the language model when topics are first extracted from transcripts. The model selects from a defined set of categories, falling back to uncategorised if none match.

Known failure modes. Multi-category stories (e.g. a political scandal about water companies) must be forced into a single primary category, losing secondary context; niche or edge topics can be misclassified.

Minimum sample size. A single story category is a simple classification; category shares and channel agenda breakdowns require a full broadcast or a week of broadcasts to represent real emphasis.

Broadcast stills

What is measured. Representative video frames captured from news broadcasts to illustrate stories, show visual framing, and verify transcript contents.

How it's calculated. Video frames are sampled upstream at a fixed interval of every 60 seconds (as declared on the copyright page). We select the frame nearest to the midpoint of the story's coverage window to represent it, using the images strictly under the fair dealing exception for criticism and review.

Known failure modes. Sampling at a fixed interval can miss brief but significant visual moments, or accidentally capture commercials, channel graphics, or presenters during transition segments.

Minimum sample size. Not a metric; stills are purely illustrative and have no statistical threshold or sample size requirements.

News density

What is measured. The proportion of a broadcast slot or a channel's total airtime spent on identified news stories rather than commercials, presenter handovers, or transitions.

How it's calculated. For a single broadcast, it is calculated as the duration of the union of all story time-windows divided by the total broadcast duration. For a channel, it is the sum of all its broadcasts' news seconds divided by the sum of their total durations, preventing short broadcasts from dominating.

Known failure modes. Stories the pipeline fails to identify read as "not news"; overlapping story windows are combined (unioned) rather than double-counted, but mis-timed story windows will skew the ratio.

Minimum sample size. A single broadcast's density is highly dependent on formatting; channel comparisons require a large sample size of at least 10 broadcasts to represent the true news-to-filler ratio.

Signature vocabulary

What is measured. The distinctive words and phrases a channel uses unusually often compared to all other channels in the archive.

How it's calculated. Calculated deterministically (without an LLM) using a log-odds ratio with an informative Dirichlet prior (alpha0 = 500). Frequencies of unigrams and bigrams are compared against the all-channel baseline. Terms must appear at least 5 times, have a statistically significant z-score of at least 1.96, and are filtered through a stopword list and a two-layer blocklist. Surfaced terms link directly to the transcript segments as receipts.

Known failure modes. Presenter names, recurring programme titles, and transcription errors can pass the filters and score highly; channels with smaller transcript volumes can have highly inflated scores. The comparison is against every other channel combined, so the channels contributing the most transcript volume are effectively part of the baseline they are measured against and will rarely show distinctive terms however characteristic their language is. An empty result means "close to the archive average", not "no distinctive voice", and the section names the channels it has nothing to show for rather than omitting them silently.

Minimum sample size. To prevent noise, a channel must have accumulated at least 5,000 total tokens in its transcripts to qualify, and terms with fewer than 5 occurrences are omitted.

Story tone: how bleak, how hard

What is measured. Two things about the stories a channel chose to run: what share of its airtime went to bad-news stories, and what share went to light ones. This is a description of story selection. It is not a measure of how a channel reported anything, and it is not a bias or fairness score.

How it's calculated. Each story is scored once by a language model on two independent axes, from its headline and summary only. No transcript is read. Valence runs from -2 (death, disaster or serious harm) through 0 (procedural or genuinely two-sided) to +2 (a clear, substantial improvement), and describes the outcome for the people in the story. Register runs from -2 (pure human interest, no consequence beyond those involved) through 0 (routine news of consequence) to +2 (war, statecraft, national economy). A channel's position is then the airtime-weighted share of its minutes spent on stories scoring -1 or below on each axis, so eight minutes on a rescue counts for more than twenty seconds on a record-breaking water lily. Positions are drawn against the airtime-weighted average of all channels rather than the middle of the scale, so a channel reads as lighter or bleaker than the average UK bulletin, not in the abstract.

Why the story and not the broadcast. Scoring transcripts would mean scoring whoever happened to be speaking. There is no speaker labelling in the archive, so a presenter reading a script, a guest arguing, and a clipped politician are indistinguishable, and a channel that books guests to argue would score as whatever its guests said. Scoring the story instead means the same event carries the same score on every channel, and the differences between channels come only from what they chose to run and how long they ran it.

Known failure modes. Stories that are good for one group and bad for another, such as a policy that improves air quality but costs drivers, have no correct valence and should land near 0; if they scatter instead, the axis is noisier than it appears. Good news is genuinely scarce, so the good-news column is small for every channel and the axis with real range is the bleak one. Unmerged duplicates of the same event inflate whichever channels aired them.

Minimum sample size. A day counts only if at least 90% of its story airtime has been scored, and the chart appears only when at least 21 of the last 30 days qualify; below that both surfaces hide themselves rather than compare channels on different denominators. Coverage is stated on the chart. A channel needs at least 10 scored stories in the qualifying days to be plotted, and channels below that are named rather than silently omitted.

Weekly overview & superlatives

What is measured. Weekly awards (Most Divergent, Closest Pair, Biggest Blind Spot, Most Contested Quote, and Vanished Fastest) and glance stats that summarize broadcaster focus over a completed week.

How it's calculated. Rendered only when the week is complete (every captured day status is final, latest data on/after the week's Sunday, and at least three data days are present). Divergence and closeness are computed via cosine similarity over weekly topic-airtime-share vectors. Blind spots require at least three channels covering a topic for at least 10 minutes (600 seconds) total. Vanished fastest requires a story peak of at least 5 minutes (300 seconds) between Monday and Friday followed by silence on at least two subsequent data days. Accused or winning channels must have broadcasted on at least three days that week.

Known failure modes. A channel that broadcasts just under the three-day qualification threshold will be omitted from awards; gaps in data capture will shrink the evidence base, though the site displays missing days as inactive chips.

Minimum sample size. The minimum qualifications (at least three days of data, completed week status, and channel eligibility thresholds) serve as the minimum sample gate.

Web article context & source spectrum

What is measured. External web press articles attached to major stories to provide context, along with a per-source editorial spectrum label indicating the outlet's political positioning.

How it's calculated. An entity-first search query built from the story title is run against a full-text index of the articles we hold. Each result carries a relevance score between 0 and 1; we keep articles at or above 0.45, and rescue those above 0.15 that also mention a key entity from the story in their own title. Political spectrum labels (left, centre-left, centre, centre-right, right) are editorial assignments we maintain per source, not per article.

Known failure modes. Relevance filtering can occasionally exclude valid articles or admit irrelevant ones with similar titles; political spectrum labels are shorthand tags for the outlet itself rather than a direct measurement of the specific article.

Minimum sample size. Not a metric; articles are context and have no statistical threshold or minimum sample size.

Trending & new story rails

What is measured. The stories surfaced on the homepage rails, highlighting trending coverage and newly detected developments.

How it's calculated. Trending scores are computed by summing a story's daily airtime with an exponential decay of 0.7 per day. A story must have aired within 24 hours of the newest data to qualify as trending. A "NEW" badge is assigned to stories first detected within 48 hours of the newest data. All windows are anchored to the newest broadcast date in the archive, not the wall clock.

Known failure modes. With sparse or missing data, the decay calculation can let a single long broadcast dominate the trending rankings; "new" status refers to the first time the pipeline processed the story, not its first broadcast globally.

Minimum sample size. Rankings are only as fresh as the last pipeline run and depend on having at least one recent broadcast within the qualification window.

Lead story comparison

What is measured. The story each channel led with — and where channels agreed or split — shown on date pages and the homepage as “What led the news”.

How it's calculated. When a channel has several analysed broadcasts in a day, the longest one counts as its flagship. Each story's time window is intersected with the flagship's opening ten minutes (600 seconds), and the story with the greatest overlap is the lead; if nothing overlaps the opening window, the earliest story is used instead and labelled “earliest story”. The consensus lead is the story that led on the most channels.

Known failure modes. Headline round-ups or teaser segments inside the opening window can outweigh the true lead; mis-timed story windows shift the overlap; when a broadcast's end time is missing, its duration is approximated from summed story airtime, which can occasionally pick the wrong flagship.

Minimum sample size. One day's lead is a single editorial decision — patterns need a run of days. The comparison renders only when at least two channels have an analysable flagship broadcast.

Broadcast Prism summaries are cross-channel consensus, not an objective account. Topics, airtime, and gaps come from automated transcript analysis and may contain errors. Treat every metric here as a map of where attention went, not a verdict on what happened.

About the summaries

Broadcast Prism summaries are shared cross-channel summaries, not objective accounts. A shared summary can be wrong, or omit what only one channel covered. Topics and coverage are derived from automated transcript analysis and may contain errors.

About Methodology Privacy Terms Copyright & sources Contact RSS (Digest) RSS (Stories)
Broadcast Still

Search across all channels and broadcasts

No results found for

Search is unavailable right now. Please try again shortly.

↑↓ to navigate ↵ to select ESC to close
Open in advanced search →