Kagi AI-On vs AI-Off: A Reproducible Search Test Protocol

Compare Kagi AI-on and AI-off search through paired queries, fixed account settings, randomized order, separate result layers, transparent limitations and a reusable ledger.

Sonar and a researcher run identical query cards through AI-on and AI-off lanes while account, query and time remain fixed.

Direct answer: Kagi documents a Search setting that can disable access to AI features. To learn what that changes for a real search job, run paired AI-on and AI-off observations while holding the account, query, lenses, domain rules, locale, time window and result depth as constant as possible.

The switch confirms a controllable treatment; it does not establish which mode is better or whether every query changes. SearchEngineAnswer has not run the sample described here, so this article publishes a reproducible protocol and ledger rather than invented results.

Confirm the control and its scope

Kagi’s July 2, 2026 changelog says it added an option under Search AI settings to completely turn off AI-based features in Search. The current settings documentation describes separate controls for AI features. Save the settings page, account configuration and access date because the available controls and their names can change.

The changelog also said Kagi planned to add the option to onboarding. A plan is not evidence that a particular onboarding flow has shipped. Test only the interface available to the account used in the study.

Choose one falsifiable question

“Is AI search better?” is too broad. Choose one primary question before building the sample: does the setting change visible modules, ordered organic URLs, source visibility, time to a usable result, or the rate at which an assessor completes a declared task?

A measure must match the research question
QuestionPrimary measureDoes not answer
Does layout change?Module presence and footprintRetrieval quality
Do visible sources change?URL and domain overlapTask success
Is evidence easier to trace?Claim-to-source verificationLatency
Does a user finish the task?Predefined completion and error rateMarket-wide preference
Is the response faster?Client-observed timing under controlsAnswer accuracy

Freeze the account profile

Use one dedicated study account and preserve a safe configuration snapshot. Record lenses, upranks, downranks, blocked domains, language, region, Safe Search, result density and any other setting that can affect retrieval or presentation. Keep credentials and personal identifiers out of the public manifest.

If the account configuration changes during collection, start a new run version or repeat the affected pairs. Labelling two unlike profiles as one condition creates a false comparison.

Build a stratified query sample

Select query types that match the intended use: navigational, current-fact, explanatory, comparison, local, shopping, troubleshooting and ambiguous questions where relevant. Define inclusion rules before viewing results. Version the exact strings and calculate a hash so a later run cannot quietly substitute easier prompts.

A small sample can show reproducible differences inside that sample. It cannot establish permanent Kagi behavior or quality for all search users.

Run paired observations

  1. Confirm the saved account profile and clear only the state the protocol says to clear.
  2. Randomize whether the pair starts with AI on or AI off.
  3. Run the exact query and save the complete result page where permitted.
  4. Switch only the documented AI control.
  5. Repeat within a declared time window.
  6. Record errors, zero-result pages and unexpected modules.
  7. Restore the starting setting and verify the next pair.

Alternating the starting condition reduces a simple order effect. It does not freeze Kagi’s live index, network conditions or concurrent product experiments, so record the time of both observations.

Hold and record the confounders

Variables to freeze or disclose
VariableControl methodIf it changes
Account profileSaved configuration versionRestart or label a new run
QueryExact versioned stringDo not treat as the same pair
Lenses and domain rulesFreeze settingsExclude or repeat
EnvironmentRecord locale, device and browserReport the difference
TimeDeclare maximum pair intervalFlag index drift risk
AI settingChange only the documented switchThe comparison is invalid

Capture the whole result, not a winner

Save the complete page or a permitted structured export, not a crop of the preferred answer. Classify all visible modules, then extract organic URLs with displayed and resolved forms. Preserve duplicate URLs, missing modules, errors and results below the first screen.

A screenshot supports auditability but is not enough for URL overlap or timing analysis. Pair it with structured observations and a content hash when possible.

Score five layers separately

Do not hide unlike outcomes in one score
LayerExample observationsTypical unit
PresentationModules, summaries and visual footprintPresence or pixels
RetrievalURL overlap, order, unique domains, freshnessCount or rank metric
EvidenceSource visibility and claim traceabilityVerified claims
TaskCompletion, time, errors and confidenceRate or duration
PerformanceClient-observed response timeMilliseconds

If a combined score is essential, publish the weights before evaluating outputs. A quieter interface can coexist with similar retrieval, while a convenient summary can make source verification harder.

Normalize URLs without erasing evidence

Keep the displayed URL, resolved URL and normalization rule. Remove a fragment or known tracking parameter only under a published rule; do not merge pages merely because they share a domain. Report both exact-URL and domain overlap when each is relevant.

For ordered top-k results, record rank as well as presence. Two conditions can share the same URLs while placing them in materially different positions.

Blind human assessment where possible

Define task success and evidence requirements before collection. Remove obvious mode labels from assessment copies when the interface and study ethics allow it, randomize presentation order and keep the person scoring task completion separate from the person who configured the switch.

Record disagreement and adjudication. Confidence should not substitute for a verified source or a completed reader task.

Download the paired-run ledger

The Kagi AI-on/off paired-run ledger records the query version, randomized order, account profile, captures, URL overlap, task outcome and exclusions. Remove the example row before use.

Keep private captures and account settings in controlled storage. The public artifact should include enough configuration to reproduce the method without exposing authentication data.

Analyze pairs before averages

Validate that each row is a true pair, then calculate per-query differences. Report the distribution, missing pairs and exclusions before a mean. One dramatic query should not stand in for the sample. With a small sample, emphasize effect sizes and uncertainty rather than a binary declaration.

Record aborted pairs separately so missing observations do not silently favor either condition.

Repeat a subset on another date to distinguish a stable mode difference from ordinary live-search variation. Label the repeat as a new run, not an invisible extension.

Publish the manifest and limitations

Publish the query set where permissible, configuration fields, rubric, exclusions, raw or redacted captures, calculation code or formula, and observation dates. State that the study covers one account profile, locale, time window and product version. Keep the Kagi reproducibility guide beside the run and use the tool-score evaluation guide to expose any weights.

Do not call the AI-off condition “no algorithms” or “unpersonalized” unless those separate properties were controlled and documented. The switch concerns AI-based features, not every system involved in search.

Primary documentation

Source review: August 31, 2026. This protocol reports no comparison result until a declared sample is run.

Keep learning

Continue this topic

Community discussion

Discuss: Kagi AI-On vs AI-Off: A Reproducible Search Test Protocol

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.