Parts of BotScope’s scoring have always been rules: phrase lists that spot an AI saying it doesn’t know you, keyword checks that confirm it described the right company, word-by-word sentiment, and a parser that reads ranked lists. Rules are fast and cheap enough to run on every answer we collect. Their weakness is that they fail quietly when AI answers change shape — and AI answers change shape all the time.

We recently audited those rules against real answers, read the disagreements by hand, and found three places where they were wrong often enough to matter. All three now judge the answer in context instead. This post explains what changed, how we tested it, and what you will see in your numbers.

TL;DR

  • Recognition (L1). Answers where the AI politely says it can’t tell which firm you mean no longer count as recognising you — the old rules could score those 100. Answers that describe a different organisation with your name now count as the wrong company.
  • Sentiment. Every mention is judged from the passage around it, not word by word. Practice-area words (“injury”, “claims”, “criminal”) no longer read as criticism, non-answers and plain list entries no longer read as praise — and praise for a different organisation that shares your name no longer counts for you.
  • Rankings (L3). Comparison tables, local business listings and headed lists are now read — tables alone are about a fifth of ranking answers, and were previously ignored. Your rank is counted among real recommendations only, and “Best for”, “Map data is currently unavailable” and similar page furniture are gone from your competitor lists.
  • Sentiment and ranking history has been re-scored with the same method, so those trend lines compare like with like. The recognition change applies from new scans (its effect on past scores was under half a point). Reports and share links you have already sent keep the numbers they had.
  • Nothing to do, and no extra credits. It applies to every project automatically.
Same answer, old scoring vs new · choose an area

L1 · does the AI know who you are?

I can’t reliably say which “Hartwell & Co Solicitors” you mean — several firms use that name. If you share the city or the website, I can tell you more.
Before · rules Recognised, right company · L1 100
Now · judged in context Did not recognise you · L1 0

The old rules looked for phrases like “I don’t have information”, and today’s models decline more politely. Then the category word “solicitors” — sitting inside the firm’s own name — was taken as proof the answer described the right company.

Hartwell & Co is an American family-owned furniture maker founded in 1952, best known for solid oak dining tables.
Before · rules Recognised, right company · L1 100
Now · judged in context Described a different organisation · L1 30

A confident, fluent description — of someone else. The old check only looked for a category word, and “family-owned” contains “family”. The answer is now compared with who you actually are.

Recognition: does the AI actually know who you are?

L1 asks the most basic question in AI visibility: when someone asks an AI about you by name, does it know who you are? The old rules answered it in two steps — look for a “don’t know” phrase, then check the answer for one of your category words or a few words from your brand description.

Both steps had blind spots. Current models rarely say “I don’t have information about…” any more; they say “I can’t reliably tell which firm you mean — several use that name.” None of the phrases matched. Then the category check passed because the category word was inside your own name — “solicitors” in “Hartwell & Co Solicitors” — so an answer that merely repeated your name counted as describing the right company. The result, in the cases where both steps misfired, was an “I don’t know which firm you mean” scored as full recognition.

Now each answer is judged directly: does it describe your organisation, a different organisation that shares the name, several possible organisations, or does it not know? Those map onto the same bands as before — 100, 30, 70 and 0 — and a judgement the model isn’t confident about is scored as “recognised, can’t confirm” (70) rather than a full 100. Your stored brand description is used as the reference, and is itself checked first: an old description that doesn’t match your categories is set aside in favour of your website and categories.

For most projects the change is small. Where it isn’t, the old number was flattering you.

Sentiment: how each mention talks about you

Sentiment used a well-known word-scoring method: add up the positive and negative words in the sentence around your name. That works for reviews on social media, which is what it was built for. It works much less well for AI answers about businesses:

  • A law firm described by its practice areas — personal injury, criminal defence, negligence claims — read as negative.
  • “If budget isn’t a concern, Hartwell & Co is hard to beat” read as negative, one word at a time.
  • A firm fined by its regulator read as neutral — no emotive words.
  • An AI saying “I can’t tell which Hartwell & Co you mean… I’d recommend checking their website” read as positive, because of “recommend”.

Every mention is now judged from the paragraph or list item it appears in: positive, neutral, negative or mixed, towards the brand it names. Mentions of your own brand get one extra check — is this passage actually about your organisation, or its parent company or group? — so praise for a restaurant, studio or engineering firm that happens to share your name no longer counts for you.

On real mentions, the old method and the new judgement agreed only between half and two-thirds of the time. We read the disagreements by hand; the new judgement was right in nearly every case. The biggest shifts are fewer false negatives from sector vocabulary and fewer false positives from non-answers and plain list entries — which also means your sentiment index may read differently, usually higher, because confident praise now registers as praise rather than being diluted by word counts.

If you had turned on YMYL mode or added custom neutral terms to compensate for the old method, you no longer need them: they now apply only to the rules-based fallback we use if the model is ever unavailable.

Rankings: where you rank when AI recommends someone

L3 is the heaviest-weighted part of your visibility score, at 35% of the overall: when someone asks an AI for the best firms, products or providers in your category, are you in the answer, and where? The old parser read numbered lists, bullets and bold lines. AI answers have moved on:

  • Comparison tables — about a fifth of all ranking answers, and more than a third of ChatGPT’s — weren’t read at all. A brand in row 1 of a table counted as merely “mentioned”.
  • Local listings — ChatGPT’s map cards and the business listings inside Google’s AI Overviews — were misread, and a “Map data is currently unavailable” placeholder regularly became the “#1 competitor”, pushing every real firm down a place.
  • Section labels — “Best for”, “Overview”, “Key strengths” — were counted as list entries and as competitors.
  • Names had to match exactly, so “Hartwell & Co Solicitors LLP” or a linked name could miss, and names were cut at hyphens (“Rolls-Royce” became “Rolls”).

Rankings now work in two halves. Code reads every list-like structure in the answer — tables, numbered and bulleted lists, headings, map cards and local listings — and cleans each entry’s name. Then each entry is judged: is it a real recommendation (a firm, brand, product or provider) or page furniture? Is it you — including your aliases, your parent company and your own products? And where an answer has more than one list, which one is the main recommendation? Your rank is counted among the real recommendations in the main list. If you appear only in a secondary list, you count as mentioned; if the answer recommends you only in prose, you are ranked by the order it names you.

We checked this against hand-labelled samples of real ranking answers, including a sample we did not look at until the method was finished. The old parser put the brand in the right score band 70–79% of the time; the new method 97%. It got the exact position right 31–45% of the time; the new method 91–100%. Competitor lists went from an average of about four junk entries per answer to 99–100% real organisations.

For most projects this means L3 goes up, sometimes substantially: being listed in a table or a local listing used to count as unranked.

How it works

The judging is done by Jev, a decision model from TypeSafe. Unlike a chatbot, it doesn’t write text: it answers narrow, typed questions — pick one of these options, yes or no — and reports how confident it is. That suits scoring. Code does everything that should be exact (reading the structure of an answer, counting ranks, applying score bands), and the model makes only the judgement calls a person would make reading the answer: is this about the same company? is “Best for” a firm?

A few things we built in, because scores that move without explanation are worse than no scores:

  • Pinned versions. The model version is fixed, so your scores don’t shift when the provider updates it; we move deliberately and say so.
  • Every judgement is stored alongside the answer it was made on, so we can audit and re-tune.
  • Rules remain as the fallback. If the model is unavailable, that answer is scored with the rules and the event is logged.
  • Your data. The model sees the AI answers we collect for you plus your brand name, aliases and description. TypeSafe is listed as a sub-processor in our privacy policy, and states that it does not train on customer requests or responses.

What changes in your numbers

  • Sentiment and rankings are re-scored across your history. The dashboard, API and MCP use the new method back to your first scan, so a jump in those charts means your visibility changed — not our scoring. Recognition changes apply from new scans onwards; for past scans the difference was under half a point overall. PDFs, share links and client reports you have already sent keep the numbers they had when you sent them.
  • Rankings and overall scores: most projects move up, some by several points overall, driven by L3.
  • Sentiment: the mix shifts — typically fewer negatives and a higher sentiment index — and the change is consistent across your history.
  • Recognition: a small number of answers score lower where the AI didn’t actually know you.
  • No action and no credits. It uses answers you have already collected.

The honest limitations

The model can still be wrong. The hardest case is a same-name business in the same line of work — another “Hartwell & Co Solicitors” in a different city. We score low-confidence recognition conservatively, and we will keep testing against real answers.

Some AI Overviews bury names in running text. When there is no list to read, you are ranked by the order the answer names you, which is a reasonable proxy but not a ranking the AI stated.

“Mentioned but not recommended” still counts for something. Being named in a ranking answer — in passing, as a caveat, or in a secondary list — still scores 20, as before: the AI knew to bring you up.

English is where the model is strongest. Answers in other languages are handled, with lower accuracy.

FAQs

Do I need to do anything? No. Every project uses the new scoring automatically, and sentiment and ranking history has already been re-scored.

Does this cost extra credits? No. It runs on answers you have already collected.

Why did my overall score go up? Most likely rankings: if AI answers list you in comparison tables or local listings, those placements now count. Open the Rankings view to see individual answers and the position we now read.

Why did a score go down? The usual reasons are an AI answer that didn’t really recognise you (recognition), or a local listing that put you lower than the old mention-order guess did (rankings). If a specific answer looks wrong to you, reply to our email or get in touch — every judgement is stored, so we can look at exactly what happened.

Can I still see my old numbers? Reports, PDFs and share links you created before the change keep their original figures. The dashboard shows the re-scored sentiment and ranking history.


Scores are only useful if you can trust what moves them. These changes make three of ours measure what they say they measure — and because history moved with them, the trend you see from here on is a trend in your visibility, not in our method.