AI web scrapers make up data when a field is missing, and the made-up value is usually real text copied from the wrong line of the page. A web extraction honesty benchmark published on September 27, 2026 tested 16 models and 3 paid extraction APIs on pages where the answer had been removed. Adding one sentence, "Use null for any field whose value is not on the page. Do not guess.", cut made-up fields from 70.7% (405 of 573) to 20.2% (116 of 574).
We rebuilt a small version of that test on a CPU with two local models, Qwen3.5 2B and 4B, to see which fixes a builder can rely on. The sentence helped the 4B model a lot and the 2B model much less. Asking for a verbatim quote and checking it in code caught none of the decoys. What worked was a second model answering yes or no for each value: with the 4B model as checker, 0 of 12 missing fields came back invented for either extractor, and none of the 115 correct values was thrown away. Below is the test, the results, and the pipeline to copy.
What the "Do Not Guess" Benchmark Found
The benchmark comes from Earn an Honest Dollar, a marketplace where providers "Offer work you perform or software you operate", so it has an interest in buyers comparing extraction services. Its design is still worth copying. It built 42 pairs of twin pages across 7 page types. Each pair differs by one row: one page shows the answer, the other does not. Both pages carry a decoy, such as an old price ("Was $493.00"), a wrong author ("Fact-checked by Omar Tamm") or an old timestamp.
Every model made up more without the sentence. Gemini 3.8 Flash went from 14 of 36 missing fields invented to 1 of 36, and GLM 5.3 from 18 of 36 to 1 of 35. Qwen 3.8 27B went from 30 of 36 to 7 of 36. Gemma 4 31B still invented 13 of 36 with the sentence. The three paid APIs were tested with the instruction only: ScrapeGraphAI invented 7 of 31, ScrapingBee 16 of 36, and Firecrawl 24 of 36. The page notes that all 24 Firecrawl answers "copied the decoy".
The same page reports a cheap fix. Used as a checker, GPT-6 Luna caught 38 of 49 made-up values without rejecting a single correct one, and checking every response cost $0.0049. The authors list their own limits: one run per contestant, synthetic pages, and paid APIs tested on free tiers.

How We Tested It on a CPU
We wanted to know whether the pattern holds for small local models, and which part of the fix does the work. The setup:
- Pages: 12 twin pairs we wrote across 6 page types (product, article, event, software release, SaaS pricing, freelance brief), 24 pages, 3 fields each. Every name, product and price is fictional. Page B is page A with one row deleted.
- Decoys: one per pair, of the kind real pages carry: a "Was" price, a photo credit or quoted source, last year's event date, an older version number in an upgrade note, a previous plan price, a "typical" budget.
- Runtime: llama.cpp
llama-serverbuild b11235 on 6 CPU threads, temperature 0, output forced to a JSON schema in which every field may be null. - Models: Qwen3.5 2B and Qwen3.5 4B, both Q4_K_M, thinking turned off.
- Conditions: baseline; baseline plus the benchmark's sentence, word for word; the sentence plus a verbatim quote for every field.
That gives 12 missing fields per model per condition, in a single run: read the numbers as direction, not a leaderboard.
What We Found
| Step (12 missing fields) | Qwen3.5 2B invented | Qwen3.5 4B invented |
|---|---|---|
| Baseline | 12 (11 copied a decoy) | 10 (9 copied a decoy) |
| + "Do not guess" | 7 (5 decoys) | 2 (2 decoys) |
| + verbatim quote | 7 (6 decoys) | 1 (1 decoy) |
| + code check of the quote | 7 | 1 |
| + Qwen3.5 2B as checker | 3 | 0 |
| + Qwen3.5 4B as checker | 0 | 0 |
When the answer was on the page, both models found it: 12 of 12 in every prompt condition. The problem is only the empty case, and in the empty case the models rarely invented from nothing. At baseline the 4B model returned "Was $329.00" as the current price, a photo credit as the author, last year's date as the event date, and "3.9.4" from the line "Upgrading from 3.9.4 or earlier?" as the latest version.
The sentence did most of the work for the 4B model, from 10 down to 2. For the 2B model it was weaker than it looks: 5 of its "fixes" came back as the word "null" inside a string, not a JSON null. A pipeline that only tests for a real null would have stored the text "null" in 5 price, author and budget fields.

Why Quote Checks Miss Decoys
A common piece of advice is to make the model cite its source and then verify the citation. We tried the strictest cheap version: the quote must appear on the page character for character, and the value must appear inside the quote. It caught 2 junk values where the 2B model had copied the field description, "software name", instead of the software's name. It caught 0 of the 7 decoys.
The reason is simple once you see it. A decoy is real text. "Was $329.00" is on the page, the quote is exact, and the value sits inside it. String checks prove a value exists, not that it answers the question. The check also cost something: it rejected 3 correct values, 2 where the 2B model left its quotes empty and 1 where the 4B model quoted the wrong line for a correct product name. Asking for quotes also roughly doubled the time per page, from 1.8 to 3.4 seconds for the 2B model and from 4.2 to 7.7 seconds for the 4B model.
Judging meaning takes a model. With the 4B model as checker, every one of the 2B model's 9 wrong values (6 decoys and 3 others) was rejected, and all 58 of its correct values were kept. The 2B model was a poor checker: it caught 4 of those 9 and threw away 2 correct values. That matches the benchmark, where the checker was a capable model, and it gives a rule of thumb: the checker should be at least as strong as the extractor.
How to Build a Null-First Extraction Pipeline
This takes an afternoon to wire into an existing scraper. It works the same with a hosted API that supports schemas, such as Gemini structured output, or with a local server.
Step 1: Make every field nullable and enforce the schema
Type each field as ["string", "null"] (see the JSON Schema null type) and send it as a schema, not as a description in the prompt. In llama-server that is response_format with type: json_schema. If null is not a legal output, the model has to put something in the field.
Step 2: Add the sentence, word for word
Put "Use null for any field whose value is not on the page. Do not guess." at the end of your instructions. Describe fields by meaning, for example "current selling price", not just "price", which gives the model a reason to skip the "Was" line.
Step 3: Treat null look-alikes as null
Before you store anything, convert "null", "none" and empty strings to a real null. Our 2B model produced the string "null" 5 times and the 4B model returned one empty string.
def clean(v):
if v is None or str(v).strip().lower() in ("", "null", "none", "n/a"):
return None
return vStep 4: Run a yes or no checker on every non-null value
Send the page, the field meaning and the value to a second call, and accept only "yes". The core of the wording we used, with the page text after it:
Field: current selling price
Value: $329.00
Quoted from the page: Was $329.00
Answer yes only if the page states this value as the current selling price itself.
Answer no if the value is not on the page, or if it is on the page but refers to
something else. Answer with one word: yes or no.Force the answer with a two-value enum, and turn every "no" into null plus a log line you can review. Use a checker at least as strong as the extractor.
Step 5: Keep a quote only for audit
Quotes help a human review the "no" log, and the code check catches copy-the-label junk, but a passed quote check is not proof. If speed matters more than audit, drop the quote and keep the checker.
Step 6: Build twin pages from your own targets
Save five real pages you scrape, delete the row that holds each field you care about, and run your pipeline on both versions. A field that comes back filled on the edited page is a decoy problem on that site, and you now know which line causes it.

Troubleshooting
- The local server dies halfway through a batch.
llama-serverkeeps a prompt cache in RAM, 8,192 MiB by default according to the server documentation. On our 12 GB machine the 4B run was killed by the out-of-memory killer twice at the same point. Adding--cache-ram 0fixed it; we also ran a single slot (-np 1). - Reasoning text you did not ask for. Qwen3.5 has a thinking mode. For extraction we switched it off by sending
chat_template_kwargswithenable_thinkingset to false, as the server documentation shows. - The checker returns broken JSON. We first capped it at 5 tokens and the answer was cut off mid-JSON. About 30 tokens is enough.
- A field holds its own label. The 2B model once wrote "software name" as the software's name. The quote check catches this; so does a larger model.
- The checker throws away good values. That is a weak checker. Our 2B checker rejected 2 of 58 correct values; the 4B checker rejected none.
What This Means for Builders
If you build price trackers, tool directories, event calendars or research sheets with AI, the dangerous output is not an obvious hallucination. It is a plausible, stale value that sits next to the one you wanted: last year's date, the old price, the photographer's name. It looks right in a spreadsheet until a reader notices.
The benchmark suggests picking an extraction service by measured results, and the spread supports that: from 1 of 36 to 24 of 36 invented with the same instruction. But even the best one will meet a page it has never seen, so the checker is the part that protects you. On a CPU it is one extra short call per filled field.
Key Takeaways
- One sentence, "Use null for any field whose value is not on the page. Do not guess.", cut made-up fields from 70.7% to 20.2% across 16 models in the benchmark.
- Most invented values are decoys copied from the page, so quote-matching in code does not catch them: 0 of 7 in our test.
- A yes or no checker at least as strong as the extractor fixed it: our 4B checker caught every wrong value and rejected no correct ones.
- Small models answer "null" as a string; clean those before storing.
What to Watch
The benchmark ran once on synthetic pages, and ours is smaller still, so the next useful number is a repeat on real pages with several runs per model. Watch whether extraction APIs start publishing their made-up rate on missing fields next to accuracy, and whether they add a built-in checker. Until then, the number that matters is the one you measure on your own twin pages.
Frequently asked questions
Why do AI scrapers make up data?
Asked to fill a field, a model prefers a plausible value to an empty one. Most invented values in the benchmark and in our test were real page text answering a different question, such as an old price.
Does telling the model "do not guess" work?
It helps. In the benchmark, made-up fields fell from 70.7% to 20.2%; our Qwen3.5 4B model went from 10 of 12 to 2 of 12. Smaller models improve less, and some return "null" as text.
Is asking for a source quote enough to stop hallucinated fields?
No. A quote check proves the value is on the page, not that it is the right value. It caught 0 of 7 decoys in our test, though it does catch junk such as a copied field label.
What is a checker model in web extraction?
A second call that reads the page and one extracted value and answers yes or no. Values that get a "no" become null. In the benchmark GPT-6 Luna caught 38 of 49 made-up values for $0.0049 in total; in our test a local 4B model caught all 9 wrong values from the 2B extractor.
Can I run this on a laptop without a GPU?
Yes. Our whole test ran on 6 CPU threads with llama.cpp. The 4B model took about 4 seconds per page to extract without quotes, and the checker adds one short call per filled field.