About the evidence framework

AI visibility is noisy. Ternith measures it that way.

We built Ternith around what our research found: one run moves, providers differ, phrasing changes outcomes, and retrieval does not prove support.

Our early tests found movement in the measurement, so we built the product around repeated runs, provider comparisons, source review, and refusals to overclaim.

Plate 00 · framework

01Questions
02Provider runs
03Retrieved sources
04Source review
05Two-run compare
06Diagnosis
Run 1
Run 2
Closer markers mean stronger agreement.
01aligned

Both runs point to compellingness. Ternith can diagnose.

The evidence changed the product.

We built Ternith after the measurement work exposed the problem we were trying to solve: AI visibility moves, and a single score hides too much.

Our first question was practical: what evidence would make a diagnosis worth trusting?

The research showed that one run can move, providers can split, phrasing can change the shortlist, and retrieved pages can fail to support the recommendation.

That finding shaped the product. Ternith repeats the run, compares providers, reviews sources, and refuses to call a diagnosis when the evidence is not strong enough.

01Single number

Many tools start with one tidy result.

02Instability found

The same surface moves under measurement.

03Framework built

The product gains checks, repeats, and refusals.

04Diagnosis earned

The result ships only when evidence supports it.

What we found.

repeat

One run is fragile.

The same question can produce a different shortlist when the system answers again.

providers

Providers differ.

ChatGPT, Claude, and Gemini expose evidence differently and recommend differently.

wording

Phrasing moves outcomes.

A small wording change can move a brand into or out of the answer.

source

Retrieval is not support.

A model can fetch a page that does not support the recommendation.

limit

Some evidence stays inconclusive.

A clean label can overstate what the data can carry.

A score hides the failure mode.

Ternith keeps discoverability, compellingness, and positioning separate because each failure needs a different fix.

74

one number

Discoverability

AI does not find enough evidence about you.

Compellingness

AI finds you, but picks someone else.

Positioning

AI picks you for the wrong buyer need.

We run studies before we make claims.

Each study starts with a buyer question and a method set before the data arrives. We define the sample, set the thresholds, and decide how we will treat weak or inconclusive evidence before we see the result.

We publish the result the study earns, even when the result is inconvenient. That habit matters because the product only works if the measurement can say no.

01
Stability

If you ask again tomorrow, will AI recommend the same brands?

published
02
Cross-provider

Do ChatGPT, Claude, and Gemini recommend the same brands?

planned
03
Question wording

Does a wider set of question wordings change the answer?

planned
04
Source review

Does the retrieved page support the brand attribution?

in product
05
Mode separation

Do discoverability, compellingness, and positioning separate cleanly?

early result

Built by

Ternith is built by Daniel Cheung.

I started Ternith because AI visibility was becoming important faster than the measurement around it was becoming reliable.

I do not want this product to sell a neat number when the evidence cannot support one. Ternith repeats the run, compares the evidence, reviews the sources, and says inconclusive when that is the honest answer.

The promise

Not certainty. Not a score. A diagnosis, built from evidence.