On 1 September 2026 our own tracker reported TypelessForm's AI visibility down from 92% to 50%, and named six rivals filling the gap. Not one of them had displaced anything. The tracker had switched ChatGPT models mid-series; four of the five lost answers came from an engine running a different model than the run before.

Webappski built that tracker, uses it on client work, and publishes TypelessForm's runs with it. So this is a story about our own instrument, caught while measuring our own product — and the reason to tell it in public is not the defect. It is what a reader is supposed to do with a number like 50%. Below is every step of the chain that produced it, the control run that isolates the cause, the part of the move that was not the instrument, and what the same three questions return today.

What is a like-for-like AI visibility comparison?

A like-for-like AI visibility comparison matches four things between runs: the questions, the engines, the models behind them, and the counting rule. Change any one of the four and you have measured two different things and subtracted them.

The questions have to match word for word. "Best voice form filling tools 2026" and "top voice form automation tools" are different retrievals with different answer sets, and a basket that quietly drifts turns a trend line into an artefact.

The engine line-up has to match. A run across four engines and a run across three are not comparable, however tempting the arithmetic looks — if a provider drops out, the missing column has to be named rather than averaged away.

The model behind the engine name has to match, and this is the one almost nobody tracks. "ChatGPT" is the name of a product, not of a fixed system: the model answering underneath it is versioned, replaced, and renamed on a schedule that has nothing to do with your reporting calendar.

The counting rule has to match. Counting the brand name inside the written answer and counting it anywhere in the raw payload — including the citation URLs, where a domain can appear a dozen times — produce different numbers from identical data. Every figure on this page uses one rule: case-insensitive occurrences of the brand name in the answer text only, with URLs and metadata excluded.

Why did our AI visibility score drop 42 points?

The 1 September 2026 run printed 50% against 92% on 13 August — a fall of 42 percentage points on paper, of which about 17 are demonstrably the instrument's doing and the rest is not established. (Percentage points of presence — how many of the twelve answers name us. The tracker also reports a separate composite index on its own 0-100 scale, and we keep the two apart throughout; more on that below.) Both runs put the same three questions to the same four engines, twelve cells each, so the basket and the line-up were honest. The model behind one of those engines was not.

That "about 17" is worth stating up front, because the tidier version of this article would say the whole fall was the instrument, and the whole fall was not. Five cells were lost. The control day measured both OpenAI models side by side, so for the three ChatGPT cells we can check each one: on two of them the newer model had already failed to name us on 13 August, so their loss on 1 September is the instrument. On the third it had named us — eight times against the older model's thirty-two — so the swap demonstrably weakened that cell without closing it, and something else finished the job two and a half weeks later. Neither "the instrument" nor "not the instrument" is true of that one. The remaining two lost cells were Gemini, where the model also changed but no same-day control exists, and Perplexity, which is read by hand. Two cells of twelve is 17 percentage points. That is what the evidence carries; the other 25 points are unattributed, and unattributed is a permissible answer.

Here is the grid, cell by cell, for the two runs and for the re-measurement the following day. Each cell is one question put to one engine; "named" means the brand appears in the engine's written answer.

Engine (model)Q1 best voice form filling toolsQ2 one-shot voice filling for e-commerceQ3 multilingual voice form fillingCells
13 Aug ChatGPT (gpt-5-search-api)NamedNamedNamed3/3
13 Aug Gemini (gemini-3.5-flash)NamedNamedNot named2/3
13 Aug Claude (proxy run)NamedNamedNamed3/3
13 Aug Perplexity (hand-read)NamedNamedNamed3/3
1 Sep ChatGPT (gpt-5.4-mini)Not namedNot namedNot named0/3
1 Sep Gemini (gemini-3.7-flash)Not namedNamedNot named1/3
1 Sep Claude (proxy run)NamedNamedNamed3/3
1 Sep Perplexity (hand-read)NamedNamedNot named2/3
2 Sep ChatGPT (gpt-5.6-luna)NamedNamedNamed3/3
2 Sep Gemini (gemini-3.7-flash)NamedNamedNot named2/3
2 Sep Claude (proxy run)NamedNamedNamed3/3
2 Sep Perplexity (hand-read)NamedNamedNot named2/3

Read the model column down the page and the story is already visible. Between 13 August and 1 September, ChatGPT went from gpt-5-search-api to gpt-5.4-mini and Gemini went from gemini-3.5-flash to gemini-3.7-flash. Nobody asked for either substitution. Five cells were lost across those two dates: three on ChatGPT, one on Gemini, one on Perplexity. Four of the five sat on an engine whose model had been replaced underneath the comparison.

The fifth is the interesting one, and we will come back to it: Perplexity, on the multilingual question, is read by hand from the same app in both runs. Same surface, same wording, same method. That loss is real, it is like-for-like, and it is still there today.

Six competitors were named. Had any of them displaced us?

No — all six names came from two answers measured on an engine whose model had changed, and both answered a different product category. The 1 September report printed them without hesitation:

Where you lost ground — top one-shot voice form filling services for e-commerce · ChatGPT · who appeared instead: Talk2Forms, Speak2Fill.ai, Smart Form Automation, ServiceMark AI Form Filler

Where you lost ground — best voice form filling tools 2026 · Gemini · who appeared instead: Fulcrum Audio FastFill, VoiceFill.ai

Every link in that chain was individually sound. Cells really were lost. Those brand strings really were in the raw answers. The extractor really did rank them by frequency. The report really did assemble a competitive narrative out of it. The conclusion was still false, and it was false for a reason no single step could catch: the answers had been produced by a model that was reading the question differently.

Four of the six came from ChatGPT's e-commerce answer. Here is how gpt-5.4-mini opened that answer, verbatim from the saved response:

"If you mean tools that let shoppers or staff fill e-commerce forms by speaking once, the best options I found are mostly browser extensions and voice agents, not full standalone 'one-shot' commerce suites."

What followed was a list of Chrome Web Store extensions and one Android app for personal autofill — a person filling in their own details on somebody else's checkout. That is a different product category from a widget a site owner installs on their own form, which is what the question asks about and what TypelessForm is. The model even disqualified one of its own picks in passing, describing ServiceMark AI Form Filler as "limited to specific SERVPRO job-intake pages, so it's not general e-commerce." That brand still arrived in our report as a competitor filling our gap.

The other two came from Gemini's answer to the "best voice form filling tools" question, which organised itself around desktop dictation apps and field-operations software. Fulcrum Audio FastFill appears there under a heading for inspection and safety reporting, described for "field operations, site inspections, and safety reporting" where inspectors narrate observations that get parsed into form fields — VoiceFill.ai sits a section further down, described as being for "customer service, call centers, and sales intake" — it transcribes phone calls and voice chats in real time and populates CRM fields so an agent does not have to take notes. Useful software, both of them. Neither is the category a site owner is shopping in when they ask which voice widget to put on their booking form.

And the one loss that was like-for-like — Perplexity on the multilingual question — named no competitor at all. The report's "who appeared instead" column for that row reads none named. So the entire competitive story came from the cells that could not support one, and the single cell that could support one had nothing to say.

Can changing the model really change the answer that much?

On 13 August 2026 the same three questions went to two OpenAI models on the same day: one named TypelessForm in 3 answers of 3, the other in 1 of 3. Same brand, same wording, same afternoon, same counting rule, both saved to disk. The only variable was the model.

OpenAI model, 13 Aug 2026Q1 mentionsQ2 mentionsQ3 mentionsAnswers naming us
gpt-5-search-api327123 of 3
gpt-5.4-mini8001 of 3

The counts above are occurrences of the brand name in the answer text, URLs excluded. We labour that because the alternative rule — counting the whole payload, citation links included — inflates some cells several times over, and a table with two counting methods in adjacent rows refutes itself before anyone reads it.

Then the arc continues, and it closes the diagnosis. On OpenAI: 13 August on gpt-5-search-api, 3 of 3. On 1 September, gpt-5.4-mini, 0 of 3. On 2 September, gpt-5.6-luna, 3 of 3 again. Nothing about TypelessForm changed across those twenty days. The instrument did, twice.

Why did the tracker change models without being asked to?

Because its model discovery filtered for a -mini name before it sorted by generation, and OpenAI's newest line had dropped the size suffix. The tracker is built to run cheaply on the current generation — that is the right default when you are re-measuring weekly. The implementation applied those two rules in the wrong order.

OpenAI's 5.6 line names its members sol, terra and luna rather than mini. Filtering on the suffix first therefore discarded the entire newest generation before generation was ever considered, leaving the picker to choose the cheapest thing in what remained — gpt-5.4-mini, two generations back — while reporting a perfectly healthy discovery. Nothing errored. Nothing warned. The run completed, the report rendered, and the number was wrong in a way that looked exactly like news.

Webappski fixed it in aeo-platform 1.11.0: generation is selected first, and the cheap tier chooses inside that generation. The next run picked gpt-5.6-luna on its own.

The useful part of this is not that our picker had a defect. It is the general shape: any tool that selects a model on your behalf can select a different one than it did last month, and most of them will not tell you. If your dashboard shows a visibility line moving and cannot show you which model produced each point, the line is not measuring only your visibility.

How much of the movement was the instrument, and how much was not?

Not all of it: Gemini ran the same model, gemini-3.7-flash, on both 1 and 2 September and still moved from 1 of 3 answers to 2 of 3. Instrument held constant, number moved anyway.

That is ordinary run-to-run variance, and it is why the tracker carries a noise floor rather than treating every single-cell change as a finding. Answer engines are not deterministic: the same question, the same model and the same day can return a different list, and one cell of movement across three questions is well inside that.

So what separates variance from a real loss, when the model is identical in both cases? Whether it persists. The cell that moved here — "best voice form filling tools" on Gemini — was named on 13 August, missing on 1 September, and named again on 2 September: it went and came back inside one run. Compare the Gemini miss we reported in the previous instalment, on the multilingual question. Gemini named us there on 11 July, stopped on 13 August, and has still not named us on 1 or 2 September — three consecutive runs, the model changing underneath along the way, and the cell staying shut. That one we called a real, measured change, and we stand by it. A single transition that reverses is noise; a cell that stays shut across runs is a signal. Neither verdict can be read off one run, which is the whole reason a run is not a report.

The other honest correction runs the same direction. Of the five cells lost between 13 August and 1 September, four sat on a swapped model — but the fifth, Perplexity on the multilingual question, did not. That one is a genuine like-for-like loss, it has nothing to do with the model axis, and it did not come back on 2 September. It is still open. We have a page addressed at that exact question, on multilingual voice form filling for international websites, and it is evidently not yet doing the job on that engine.

An article that blamed the whole 42 points on the model swap would be wrong in precisely the way the 1 September report was wrong: it would take a real signal, attach a tidy single cause, and stop looking. The discipline is not scepticism about drops. It is refusing any conclusion — including a flattering one — that the cells cannot carry.

What does the number say today?

On 2 September 2026 TypelessForm was named in 10 of 12 answers, a presence score of 83%, across ChatGPT, Gemini, Claude and Perplexity. ChatGPT 3 of 3 on gpt-5.6-luna, Gemini 2 of 3 on gemini-3.7-flash, Claude 3 of 3 on our proxy run, Perplexity 2 of 3 read by hand. Both open cells are the same question, multilingual, on Gemini and Perplexity.

Separately, and it is a different quantity, the tracker's Unified Visibility Index moved from 69 to 83 — up 14 index points, not percentage points — between those two runs. Presence counts how many answers name you; the index also weighs how prominently and how positively you are described, so the two numbers are not interchangeable and we never put them in one table. For the avoidance of the confusion this whole article is about: on 13 August the presence score was 92% and the index that day was 87, while the numeral 92 on webappski.com refers to the index on 11 July. Same digits, three different quantities.

And the run that produced today's improvement carries its own warning label, printed by the fixed tracker itself, in the report, without being asked:

"A note on the measurement: ChatGPT ran gpt-5.4-mini last time and gpt-5.6-luna this time. Different models read the same question differently and return different lists. On ChatGPT there is no overlap at all between the two runs — not one answer was measured the same way twice, so nothing on that engine is a like-for-like comparison. So some of the movement above belongs to that change rather than to your visibility."

That is the tracker discounting its own good news. The 1 September report, before the fix, had no such paragraph — which is exactly why its confident 42-point story went out unqualified. The same release also added a qualifier to the competitor list that caused all the trouble: where the old build printed the "filling the gap" names flat, the current one appends, when the lost cells came from a changed engine, "treat this as a list to check, not a confirmed ranking."

How should you read your own AI visibility number?

Read it as a measurement with an instrument attached — a number without its date, model, engine line-up and counting rule is a figure, not a fact. Five questions settle whether two of your numbers can be subtracted at all:

  • Which model produced each point? Not which engine. "ChatGPT" is a product name; ask your tool to print the model id per cell, and if it cannot, treat every line it draws as indicative.
  • Was the question wording identical? A silently edited query is a new measurement wearing the old label.
  • Was the engine line-up identical? Three engines versus four is not a dip, it is a different denominator. If a provider is missing, name it in the same breath as the number.
  • Was the counting rule identical? Brand name in the answer text is a different quantity from brand name anywhere in the payload, and citation URLs will happily supply the difference.
  • Does a named competitor come from a comparable cell? A "who replaced you" list built from cells measured on a swapped model is a list of what that other model happens to name. Check it; do not act on it as a ranking.

None of this requires our tool. It is the reading discipline that any answer-engine number needs, and it is most of what separates an AI-visibility dashboard from an AI-visibility measurement.

How this differs from what we published before

We have written about model generations before — in April 2026 — and reached the opposite conclusion. That piece, our Day 33 report, found that a week-one ChatGPT score had been measured on a previous-generation model and concluded: "The number that counts is what the current browser shows." Pick the right instrument, and the number is true.

We no longer think that is enough, and the two comparisons are not of the same quality. The Day 33 one was confounded, as its own FAQ concedes — browser versus API, different days, different model behaviour, all moving at once. The 13 August control pair on this page moves one variable: two OpenAI models, one afternoon, one API surface, one counting rule, both responses saved to disk. That is a cleaner isolation, and it points somewhere else. There is no privileged instrument whose number is finally the real one. There is only a number with its instrument named, or a number you cannot use.

Two other pages of ours sit next to this one without overlapping it. The free, open-source AEO tracker covers the outer version of this problem — four different trackers returning four different numbers for the same brand. This is the inner version: one tracker, one brand, two numbers, three weeks apart.

And our open visibility record publishes the 33%-to-100% arc and the 13 August run at 92% presence, 11 of 12, with a public report for every point on the line. That page's series stops there by design: the 1 September and 2 September runs live on this page, because they are a story about the measurement rather than another point on the trend. If you are reconciling the two, the 13 August report is published in full; the two September runs are not yet, and we will link them here when they are.

Check it yourself

The tracker is open source, which is the only reason a self-measured number is worth publishing at all. aeo-platform is a command-line tool you can install and run against any brand, including ours, and every raw answer it saves is on your disk to inspect. Version 1.11.0 is the one that chooses generation before price and prints the measurement caveat above.

If you would rather have the reading done than do it, Webappski — the studio that builds TypelessForm and the tracker — will run the first AI-visibility audit on your site and send you the report at no cost. You will get the model ids, the cells, and the caveats, including the ones that make the number smaller.