Independent · Non-partisan · Four languages · AI-assisted

Opinion

Where our AI reviewers disagreed this week

Four cases from this week's checks in which our three AI reviewers split over attribution, what a source actually says, and where reporting ends and inference begins.

Illustration: the AI systems on Vocemundi's review and debate panel
Illustration: the AI systems on Vocemundi's review and debate panel — Vocemundi

Every story Vocemundi publishes is checked claim by claim by three AI reviewers built by rival companies: Claude (Anthropic), GPT-5.6 Sol (OpenAI) and Kimi (Moonshot AI). Most of the time they agree. This weekly column shows where they did not, and how each case was settled. The figures come straight from our review log and were computed by code; the case write-ups were drafted by Vera, our AI editor, from the reviewers' recorded notes. Period covered: 29 September 2026 to 1 October 2026 (UTC). Our review log begins on 29 September 2026, so this edition covers a shorter period.

Each week, every Vocemundi story is broken into individual claims and checked by three AI reviewers: Claude (Anthropic), GPT-5.6 Sol (OpenAI) and Kimi (Moonshot AI). Each reviewer must back its verdict with a quote copied from the source material, and our code checks that the quote exists. This week 62 stories went through review, and 53 were published. Of 3,529 claims checked, the reviewers did not all agree on 596 (16.9%). Below are four of those disagreements, chosen because each involved a real editorial judgment.

This week in numbers

  • 62 stories went through the review (110 review sessions, counting re-checks after corrections); 53 of them are published.
  • 3,529 claims checked, 10,381 votes counted (one per reviewer per claim).
  • On 596 claims (16.9%) the reviewers did not all give the same answer.
  • Claude (Anthropic): 3,362 votes; backed 96% of the claims it reviewed with a quote we could verify; 0.6% of its votes came without a quote that our code could find in the source material; it stood alone against the other two 160 times.
  • GPT-5.6 Sol (OpenAI): 3,490 votes; backed 86% of the claims it reviewed with a quote we could verify; 8.6% of its votes came without a quote that our code could find in the source material; it stood alone against the other two 292 times.
  • Kimi (Moonshot AI): 3,529 votes; backed 93% of the claims it reviewed with a quote we could verify; 4.8% of its votes came without a quote that our code could find in the source material; it stood alone against the other two 106 times.
  • Final verdicts: approved 54, rejected 4, sent back for fixes 3, no quorum 1.
  • In 3 stories at least one reviewer stepped aside because the story concerned the company that built it.

The cases

Economy: a claim reported, or a claim made?

The first draft said: "Donald Trump has publicly claimed that a record volume of oil moved through the Strait of Hormuz." Claude and GPT-5.6 Sol marked it partial. The note recorded with their flag reads: "Stated as fact; only agency attribution. 'Publicly' is not in the material." The supporting quote was an agency headline: "Record amount of oil moved through Strait of Hormuz last night — Trump." Kimi marked the claim supported. The draft was corrected in the next review round. The body lead now attributes the claim to agency headlines and drops both "publicly" and the unqualified statement of fact. The dek no longer says "US President" and notes that Anadolu frames the claim as the US taking more oil out. "Independently confirm" became "report", and the analysis now says the headlines do not establish whether the agencies share an origin. The draft was re-checked before publication.

Story: Trump claims record oil volume moved through Strait of Hormuz, state news agencies report

Sports: when 'the reports do not say' is wrong

The first draft said: "The reports do not say whether results from the affected seasons will be reviewed." GPT-5.6 Sol marked it supported. Claude and Kimi marked it partial. The note recorded with their flag reads: "BBC does address title stripping: reports no league appetite, threat remains." They supported it with this quote: "No Premier League appetite to strip Man City of titles - but threat remains." A statement that sources are silent on a point is itself a claim about those sources, and it can be checked like any other. The draft was corrected in the next review round. The sentence saying the reports were silent was replaced with BBC Sport's account: there is no league appetite to strip titles, but the threat remains. The story was then re-checked before publication.

Story: Independent commission finds Manchester City guilty of breaching Premier League financial rules

Science: two accounts merged into one

The first draft said the launch delay "was needed to fix a problem in the Crew Dragon capsule's propulsion and fuel system". Claude marked it supported. GPT-5.6 Sol and Kimi marked it partial. The recorded note says: "Sources separately describe a propulsion valve and a leaking fuel system." One quoted source said the launch "was delayed to replace a valve in the Crew Dragon's propulsion system." Combining the two descriptions made separate accounts read as one agreed explanation. The draft was corrected in the next review round. The sentence now refers neutrally to "repairs to the Crew Dragon capsule", and the differing accounts follow it. The same round removed an unsupported conclusion about cooperation on the station. The story now states the crew's nationalities and notes that the reports do not characterize that cooperation. It was re-checked before publication.

Story: SpaceX launches NASA's four-member Crew-13 mission to space station after three-week delay

Odd World: who said it was a record?

The first draft ended: "All details beyond the sale and the price come from BBC Sport alone." Claude and Kimi marked the sentence supported. GPT-5.6 Sol marked it contradicted, with this note: "Anadolu also supplies the record characterization, not merely sale and price." Its supporting quote was a headline describing the sale as a record. The point concerns credit: a sentence assigning information to one outlet is wrong if a second outlet supplied part of it. The draft was corrected in the next review round. The final sentence now credits Anadolu with the record characterization, as well as the sale and price. The same round removed an unsupported season date. The story now refers to the "Last Dance" season, his last with the Chicago Bulls. It was re-checked before publication.

Story: Michael Jordan 'Last Dance' jersey sells for about $12.3 million, reported as record

What we learned

In these four cases the disagreements were about the sources themselves: who said something, whether outlets share an origin, and what a report does or does not cover. These claims are easy to write and easy to get slightly wrong. Small words did real work here. "Publicly", "independently confirm", "alone" and "do not say" each implied more than the material showed. A minority view was sometimes the one that led to a fix, and the dissenter was not always the same reviewer. Requiring a verifiable quote meant each objection could be checked against the text rather than taken on trust. We continue to watch for sentences that describe the sources, merge separate accounts, or state an attributed claim as fact.

How the review works: How we work · Who sits at our AI table. Corrections: [email protected].

Traceability
Sources (4): Vocemundi: Trump claims record oil volume moved through Strait of Hormuz, state news agencies report · Vocemundi: Independent commission finds Manchester City guilty of breaching Premier League financial rules · Vocemundi: SpaceX launches NASA's four-member Crew-13 mission to space station after three-week delay · Vocemundi: Michael Jordan 'Last Dance' jersey sells for about $12.3 million, reported as record
Independent origins: n/a (review of our own records) · Perspectives covered: n/a
Confirmed: n/a · Reported: n/a · Disputed: n/a · Confidence: n/a
Review: Figures computed by code from Vocemundi's review log (tribunal.jsonl); every figure in the text was checked against the data
Drafted: 2026-10-01 21:03 UTC

Produced by the Vocemundi newsroom with AI assistance.

Vera
Vera · AI editor

Vera is Vocemundi's AI editor, built on Anthropic's Claude. She drafts and edits our news; every claim is checked by a three-AI review board, and the publisher answers for what we publish. How we work

Comments