We Made Four AI Models Compete to Write Our News. Here’s What Won — and Why

✍️ ACT ORIGINAL✓ HUMAN-REVIEWED⏱ 3 min read

The short version

Before letting AI write a single published word, we ran a live bake-off across four models on the same story. The results changed which one we trust — and revealed something we didn't expect.

Most “AI news” sites quietly pipe a single model into a publish button and hope for the best. We wanted to do the opposite — and do it in the open. So before our News Factory wrote a single published word, we ran a controlled test: four different AI models, the same source story, the same instructions, judged side by side. This article is the result, and it is also the first thing our system has produced about itself.

The setup

  • One real story: Google’s redesign of its search box, sourced from live reporting.
  • Four models: DeepSeek V3.2, DeepSeek V4 Flash, Google Gemini 2.5 Flash, and OpenAI GPT-4o mini.
  • Identical prompt, identical rules: summarise the source in our own words, invent nothing, structure it as a news brief.
  • Judged on four things: factual accuracy, depth, clarity of the news voice, and cost per article.

What each model did

DeepSeek V3.2 produced the strongest piece — around 400 words, well-structured, and accurate to the source without padding. It read like a real news brief: confident, specific, and free of the vague filler that gives AI writing away. It is now our primary writer.

Gemini 2.5 Flash came a close second. It was more conservative and slightly shorter, but that restraint is a feature, not a bug — it never over-claimed or reached past what the source supported. It is our fallback for exactly that reason.

GPT-4o mini was the weakest of the four. It was coherent and fast, but leaned on generic phrasing like “evolving digital landscapes” — the kind of language that sounds like content but says very little. Fine for filler; not what we want on a news page.

The result we didn’t expect

The most revealing outcome came from our cheapest model. When DeepSeek V4 Flash tried to read the source, it hit a bot-verification wall and received a security-check page instead of the article. Faced with no real information, it did not guess. It stopped and said, plainly, that the source text contained nothing to write about.

For most use cases that is a failed run. For a news engine, it is the single most important behaviour we saw all day. A model that refuses to fabricate when it cannot verify the facts is worth more than a model that writes a confident, fluent article out of thin air. It also exposed a real weakness in our own pipeline — source fetching that can be blocked — which we are now hardening. That is the whole point of testing in the open.

How we’re judging quality, honestly

We are not optimising for whichever model writes the most words, or the cheapest one on paper. Our order of priority is deliberate:

  • Accuracy first. A wrong fact on a news site costs more than any model ever saves.
  • Then depth and voice. The piece has to actually inform, in a consistent house style.
  • Then reliability. It must behave the same way on story number five hundred as on story one.
  • Cost last. At a fraction of a cent per article, the price difference between these models is real but small — not a reason to accept worse writing or weaker integrity.

What we’re shipping

Our News Factory now runs DeepSeek V3.2 as its primary writer, with Gemini 2.5 Flash as an automatic fallback, and a hard rule inherited from the test: if the source cannot be verified, the engine does not invent — it holds the story. Every article carries a visible “AI-generated, human-reviewed” label and a link to its source, and every one ends with a breakdown of how it was made.

Why it matters

Anyone can point an AI at a publish button. The harder, more honest question is which AI, judged on what, and what happens when it doesn’t know the answer. We think showing that work — the tests, the trade-offs, even the failures — is more convincing than any claim we could make about it. This is the first of many notes documenting how our systems are built. The engine that wrote most of this page is the same one we can build for you.

⚙️ How this article was made — fully automated
01📡 ScanOur engine watches trusted AI & tech sources in real time.
02🤖 WriteAI drafts an original summary in the ACT house style.
03🎨 IllustrateA custom hero image is generated for every story.
04📤 PublishReviewed, posted, and shared to social — hands-free.

This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →

Share this project

Leave a Reply