AppHeard

Blog · Written by Maya Chen · Aug 19, 2026

How AppHeard asks ChatGPT, Gemini and Perplexity the same question

Why we hit official APIs with web search on, sample the same prompt more than once, and never treat a single reply as the ranking.

Three AI answer panes side by side on a desk, no text overlay

AppHeard exists because a founder asked a simple question and got three different answers. They typed "best jewelry identifier app for iPhone" into ChatGPT, then Gemini, then Perplexity. One named them. One named a competitor and linked a listicle. One returned a paragraph with no App Store link at all. Which of those is "what AI said"?

The honest answer is: all three, and none of them if you only asked once. Consumer chat is personalized, remembered, and slightly different the next morning. A screenshot is a story. A said rate is a measurement. This post is how we turn the same discovery question into a number you can defend — and why the number is a rate, not a winner.

The question is shared. The run is not.

Every tracked category owns a set of discovery prompts: category questions, price questions, beginner questions, "for iPhone" questions. Those prompts are not written per customer. If two studios track jewelry identifier apps, they share the same questions. One run of "best free antique-identifier app for iPhone" serves everyone in that niche.

That sounds like a cost trick. It is also the only way the comparison is fair. If we let each account invent its own wording, we would be scoring who is better at prompt engineering, not who the engines name. Shared prompts also mean a competitor you have never added can still appear in the answer. The field is the field.

We do write a smaller branded set — questions that name your app — so we can check whether the answer describes you accurately. Those runs are yours. They do not leak into someone else's said rate. Discovery prompts are the ones that decide whether a buyer who has never heard of you gets pointed at your listing.

Official APIs, web search on

We do not scrape the ChatGPT box, the Gemini overlay, or the Perplexity thread. We call the official APIs for OpenAI, Google, and Perplexity, and we turn web search on for every discovery run. The methodology line on every public page is there for a reason: API responses closely approximate what a consumer sees, but they are not identical. Memory, location, and the last ten chats the user had will move the consumer surface. They should not move a measurement product.

Web search is not optional. A completion with search off is a model's prior. A completion with search on is a model's prior plus the pages it chose to read. Those pages are the work. If Perplexity cites a 2024 roundup that never added your listing, you will lose that question until the roundup changes or something newer outranks it. We store the cited domains with the answer so the next move is a URL, not a vibe.

Each engine still has its own voice. ChatGPT often writes a shortlist and hedges. Gemini is likelier to name a handful of apps and stop. Perplexity is likelier to attach citations you can open. Treating those three replies as one "AI" is how people talk. Measuring them as one engine is how people lie to themselves.

One reply is not a ranking

Ask the same prompt twice and you will not get the same paragraph. Sometimes the named apps stay and the order flips. Sometimes a linked-only mention becomes a name. Sometimes an app disappears for a day and comes back. That is not noise we ignore; it is the thing we are measuring.

So we sample. A question is not "done" when one engine has spoken. The free check asks the mini-set across all three engines. Paid tracking pulses high-demand questions daily and rotates the full three-engine matrix through the week. When a high-demand answer flips, we re-ask the same day before we treat the flip as real.

The unit that lands on the desk is not "ChatGPT said Jewelio." It is "this question, this week, this many samples, this many times your name appeared, this many times only the listing was linked." Said rate is mentioned divided by samples. Linked is a different column because a store badge without the name is a different kind of visibility. Founders who only watch the percentage miss the linked-only weeks — the weeks when the engines already know the listing and still will not say the word.

What we keep from each run

A finished run is not a screenshot. We keep:

  • The prompt text and which engine answered
  • Whether your app was named, linked, or absent
  • The other apps named beside you, when we can resolve them to a listing
  • The cited domains, classified when we can tell a listicle from a docs page
  • A short excerpt, so the desk can show the sentence instead of asserting it

We do not keep a consumer login. We do not fine-tune a model on your category. We do not invent a fourth engine and average it in. Three is the set, and the number is never rounded up on the marketing page.

The excerpt matters more than people expect. A said rate of 0% with a paragraph that describes your exact job-to-be-done is a different problem from a said rate of 0% with a paragraph about a camera app. The first is a naming problem. The second is a category problem. You cannot see that distinction in a dashboard that only stores booleans.

Why the three engines disagree

They are not looking at the same index. They do not weigh the same publishers. They do not have the same appetite for App Store links. In a jewelry or antique niche it is common to be named on Perplexity, linked on Gemini, and absent on ChatGPT in the same week. That is not a bug in the tracker. That is the market.

If we only asked the engine that already likes you, the score would climb and the installs would not. If we only asked the engine that ignores you, you would churn. The desk shows the disagreement on purpose. Engine coverage is a component of the visibility score because being named once on one surface is not the same as being the default answer.

The practical move is almost never "write a better ChatGPT prompt." You do not control that box. The move is to change the pages the engines already cite, or to become a page they are willing to cite. That is why next moves are URLs and threads, not copy tweaks inside a chat window.

Daily pulse, weekly matrix

Cost and freshness pull in opposite directions. Asking every prompt on every engine every day would burn the budget and still leave you staring at sampling noise. Asking once a month would be cheap and stale.

The compromise we ship: the highest-demand questions get a daily pulse, usually on Perplexity, because that surface moves first when a listicle changes. Every tracked question also has a weekday. On that day it runs the full matrix — ChatGPT, Gemini, Perplexity — with more than one sample. Over a week you get three-engine coverage without paying for a three-engine burst on every prompt every morning.

When a high-demand answer changes, we do not wait for the weekday. We burst the same question across all three engines that day and only then decide whether to alert. False flips are worse than a quiet day. A founder who gets a "you disappeared" email that reverses by dinner will not open the next one.

Free accounts are not on that cron. The unpaid desk is a teaser of a real measurement: one prompt in the clear, the rest locked. Daily re-asks are what the subscription buys, because daily re-asks are what the engines cost.

What this is not

It is not a rank-tracking suite for Google. It is not an ASO keyword tool with an AI badge glued on. It is not a claim that we see exactly what every iPhone user sees in the ChatGPT app after they have talked to it for a year.

It is also not a firehose of raw completions. The product is the rate, the named field, the citations, and the next page to touch. If you want to read a full answer we keep it, but the work is in the structure we pull out of it.

The scoring math — mention rate, position, engine coverage, the sentiment slot we leave pending until there is branded data — stays in sync with the code. This post is the asking, not the weighting.

Frequently Asked Questions

Do you scrape ChatGPT, Gemini or Perplexity?

No. Every discovery run goes through the official API for that engine, with web search enabled. We say so on the public pages because the consumer apps can still differ.

Why not just ask one engine?

They disagree, often on the same day. A single-engine score is a story about that vendor's index, not about whether AI names you.

Why is the result a rate instead of a yes or no?

The same prompt does not return the same paragraph twice. A rate across samples is the only number that survives a refresh.

Can I write my own prompts?

Paid plans can add a few custom prompts per app. The shared category set stays shared, so the comparison against other apps in the niche remains honest.

How is this different from googling the prompt yourself?

You can do that. You will get one reply, no citations stored, no competitor field, and no history. AppHeard is the log of those replies, asked the same way, every day.

Check the questions we would ask about you

Paste an App Store link into the free check. We write the discovery prompts from the listing, ask all three engines, and show you the first result in the clear. That is the same asking this post describes — just once, so you can see whether the work is even for you.

Written by
Maya Chen

Measurement lead. Designs how AppHeard asks ChatGPT, Gemini and Perplexity, and how those answers become a said rate.

Find out what AI said about your app

The check is free, takes about a minute, and needs no account.

More from the desk