Perplexity Photon Speeds Retrieval but Does Not Explain Citations
Perplexity's Photon infrastructure and Fast Search mode change retrieval speed and cost. Neither explains which sources become citations, so a paired benchmark is still required.
Perplexity has replaced its retrieval and ranking engine with a system called Photon. The September 24 engineering account reports a much faster internal search stage and a new fast preset for the Search API. For publishers, the useful distinction is where this system sits: it selects candidate pages before an answer is written. A faster selector does not, by itself, tell us which publisher will be cited.
That boundary is easy to miss when a search product reports latency, answer quality and citations in the same announcement. Perplexity’s technical account of Photon describes the retrieval and ranking machinery; it does not publish a before-and-after distribution of cited publisher domains, referral clicks or inclusion rates.
The step Photon actually changes
Imagine a question that could be answered by a product page, an independent review and a technical document. An index has to find plausible pages, rank a smaller set, and pass information about them to the answer system. That system may read more content, compose a response and decide which links to display. Photon is in the candidate selection and ranking stage. It is not a public rulebook for the later citation decision.
Perplexity says its previous adapted engine struggled as the index and per-document information grew. Data outside available RAM caused slow disk reads, and index merging could worsen the slowest requests. Photon uses compact posting lists for candidate retrieval and separate per-document records for ranking details. It checks a batch of needed records in cache, then overlaps disk reads for misses. Index construction happens on different nodes from live query serving, with a new serving group warmed before it receives traffic.
That design matters because a search service has a limited time budget to consider documents. It may have to choose a bounded candidate set rather than exhaustively inspect every plausible page. The architecture makes the selection stage more efficient in Perplexity’s system. It does not establish that the service now indexes more sites, favors longer pages, prefers a particular schema field, or cites more sources per answer.
Three performance numbers, three different claims
Perplexity reports that the p99 latency of its internal retrieval-and-ranking stage fell from about 800 milliseconds to 65 milliseconds after migration. The company also says its new Search API fast preset returned a single search call in 160 milliseconds at the median and 230 milliseconds at the 95th percentile. Those are not the same timer: one measures an internal stage, the other an API call. Neither measures how long a reader waits for a complete answer with citations.
The company evaluated the fast preset on six agent benchmarks spanning 3,554 selected tasks. Its task-weighted aggregate score was 64.3% at an estimated $59.73 in total model-plus-search cost, compared with 64.0% at $187.60 for the default Perplexity preset. The roughly 68% cost reduction is a vendor-run result for that task set and cost model, not a published independent quality audit. Perplexity also cautions that provider-reported search-call latency charts do not share a controlled measurement setup.
The strongest conclusion is therefore operational: Perplexity says it can serve this search stage with much lower tail latency while retaining comparable aggregate benchmark performance under its chosen evaluation. The result does not justify a publisher claim such as “Photon increased our chance of citation by 68%.” Cost per task, retrieval latency and publisher citation share have different numerators and denominators.
What a publisher can test without guessing a ranking factor
Keep an existing panel of exact questions and record the Perplexity surface, account state, locale, date, full answer, visible cited URLs and resolved destinations. Repeat each question. Compare the same panel before and after a documented change, and label any gap in the capture history. If only a post-launch sample exists, call it a baseline, not a Photon effect.
For a cited page, inspect whether the link supports the claim beside it and whether a measurable visit followed. For an uncited page, do not assume Photon rejected it. The page might not have been crawled, might have lost candidate selection, might have been read but not used, or might have informed an answer without a displayed link. Our citation-failure map separates those possibilities. Perplexity’s Q2D-Web retrieval benchmark addresses an earlier candidate-search question, not this release’s publisher traffic outcome.
A useful change log would keep four columns: source eligibility, candidate or answer evidence when observable, displayed citation, and referral. Most publishers cannot see Perplexity’s internal candidate list. Mark that field unknown rather than inferring it from the final answer. If citation mix changes, investigate alongside crawl access, question wording, model or interface changes and normal answer variation before assigning a cause.
The publisher implication is a boundary, not a tactic
Photon is evidence that Perplexity treats retrieval infrastructure as a core part of AI search. It is not evidence of a new public optimization requirement. A publisher’s durable work remains making important pages accessible, precise, current and independently useful, then measuring whether they are discovered, cited and visited as separate outcomes. The next meaningful test is a repeated citation-and-referral panel, not a sitewide rewrite to match an undocumented ranking system.
Fast Search changes the cost and latency hypothesis
Perplexity’s September 30 changelog adds search_type: "fast" to the Search API and the Agent API web_search tool. Perplexity lists Fast Search at $1 per 1,000 invocations, compared with $2.50 for standard web search, and says it is approximately 800 milliseconds faster. Agent API model-token charges remain separate.
Those numbers make Fast Search a useful benchmark candidate, not proof of equal retrieval quality. Photon describes the retrieval infrastructure. Fast Search is an exposed request mode with a price and latency claim. A buyer still needs paired results before deciding where it belongs.
- Run the same timestamped query set through fast and standard Search.
- Measure median and p95 response time rather than one demonstration.
- Compare URL overlap, domain diversity, publication freshness, and duplicate results.
- Judge whether each returned page supports the query, not merely whether a URL was returned.
- Report search invocation charges separately from model tokens and application cost.
Until that paired benchmark exists, this article treats 800 milliseconds as Perplexity’s claim and does not assume that the lower-priced mode preserves the same recall, freshness, or citation-support quality.
Keep learning
Continue this topic
Next in this topic
Cloudflare Pay-Per-Use Separates Access Fees From Content Use
Earlier in this topic
CITECHOICE Study: More Citation Credit, Uncertain First-Citation Gains
AEO & AI Search
Ask a question or join the discussion