Kagi AI-On vs AI-Off: A Reproducible Search Test Protocol
Compare Kagi AI-on and AI-off search through paired queries, fixed account settings, randomized order, separate result layers, transparent limitations and a reusable ledger.
Direct answer: Kagi documents a Search setting that can disable access to AI features. To learn what that changes for a real search job, run paired AI-on and AI-off observations while holding the account, query, lenses, domain rules, locale, time window and result depth as constant as possible.
The switch confirms a controllable treatment; it does not establish which mode is better or whether every query changes. SearchEngineAnswer has not run the sample described here, so this article publishes a reproducible protocol and ledger rather than invented results.
Confirm the control and its scope
Kagi’s July 2, 2026 changelog says it added an option under Search AI settings to completely turn off AI-based features in Search. The current settings documentation describes separate controls for AI features. Save the settings page, account configuration and access date because the available controls and their names can change.
The changelog also said Kagi planned to add the option to onboarding. A plan is not evidence that a particular onboarding flow has shipped. Test only the interface available to the account used in the study.
Choose one falsifiable question
“Is AI search better?” is too broad. Choose one primary question before building the sample: does the setting change visible modules, ordered organic URLs, source visibility, time to a usable result, or the rate at which an assessor completes a declared task?
| Question | Primary measure | Does not answer |
|---|---|---|
| Does layout change? | Module presence and footprint | Retrieval quality |
| Do visible sources change? | URL and domain overlap | Task success |
| Is evidence easier to trace? | Claim-to-source verification | Latency |
| Does a user finish the task? | Predefined completion and error rate | Market-wide preference |
| Is the response faster? | Client-observed timing under controls | Answer accuracy |
Freeze the account profile
Use one dedicated study account and preserve a safe configuration snapshot. Record lenses, upranks, downranks, blocked domains, language, region, Safe Search, result density and any other setting that can affect retrieval or presentation. Keep credentials and personal identifiers out of the public manifest.
If the account configuration changes during collection, start a new run version or repeat the affected pairs. Labelling two unlike profiles as one condition creates a false comparison.
Build a stratified query sample
Select query types that match the intended use: navigational, current-fact, explanatory, comparison, local, shopping, troubleshooting and ambiguous questions where relevant. Define inclusion rules before viewing results. Version the exact strings and calculate a hash so a later run cannot quietly substitute easier prompts.
A small sample can show reproducible differences inside that sample. It cannot establish permanent Kagi behavior or quality for all search users.
Run paired observations
- Confirm the saved account profile and clear only the state the protocol says to clear.
- Randomize whether the pair starts with AI on or AI off.
- Run the exact query and save the complete result page where permitted.
- Switch only the documented AI control.
- Repeat within a declared time window.
- Record errors, zero-result pages and unexpected modules.
- Restore the starting setting and verify the next pair.
Alternating the starting condition reduces a simple order effect. It does not freeze Kagi’s live index, network conditions or concurrent product experiments, so record the time of both observations.
Hold and record the confounders
| Variable | Control method | If it changes |
|---|---|---|
| Account profile | Saved configuration version | Restart or label a new run |
| Query | Exact versioned string | Do not treat as the same pair |
| Lenses and domain rules | Freeze settings | Exclude or repeat |
| Environment | Record locale, device and browser | Report the difference |
| Time | Declare maximum pair interval | Flag index drift risk |
| AI setting | Change only the documented switch | The comparison is invalid |
Capture the whole result, not a winner
Save the complete page or a permitted structured export, not a crop of the preferred answer. Classify all visible modules, then extract organic URLs with displayed and resolved forms. Preserve duplicate URLs, missing modules, errors and results below the first screen.
A screenshot supports auditability but is not enough for URL overlap or timing analysis. Pair it with structured observations and a content hash when possible.
Score five layers separately
| Layer | Example observations | Typical unit |
|---|---|---|
| Presentation | Modules, summaries and visual footprint | Presence or pixels |
| Retrieval | URL overlap, order, unique domains, freshness | Count or rank metric |
| Evidence | Source visibility and claim traceability | Verified claims |
| Task | Completion, time, errors and confidence | Rate or duration |
| Performance | Client-observed response time | Milliseconds |
If a combined score is essential, publish the weights before evaluating outputs. A quieter interface can coexist with similar retrieval, while a convenient summary can make source verification harder.
Normalize URLs without erasing evidence
Keep the displayed URL, resolved URL and normalization rule. Remove a fragment or known tracking parameter only under a published rule; do not merge pages merely because they share a domain. Report both exact-URL and domain overlap when each is relevant.
For ordered top-k results, record rank as well as presence. Two conditions can share the same URLs while placing them in materially different positions.
Blind human assessment where possible
Define task success and evidence requirements before collection. Remove obvious mode labels from assessment copies when the interface and study ethics allow it, randomize presentation order and keep the person scoring task completion separate from the person who configured the switch.
Record disagreement and adjudication. Confidence should not substitute for a verified source or a completed reader task.
Download the paired-run ledger
The Kagi AI-on/off paired-run ledger records the query version, randomized order, account profile, captures, URL overlap, task outcome and exclusions. Remove the example row before use.
Keep private captures and account settings in controlled storage. The public artifact should include enough configuration to reproduce the method without exposing authentication data.
Analyze pairs before averages
Validate that each row is a true pair, then calculate per-query differences. Report the distribution, missing pairs and exclusions before a mean. One dramatic query should not stand in for the sample. With a small sample, emphasize effect sizes and uncertainty rather than a binary declaration.
Record aborted pairs separately so missing observations do not silently favor either condition.
Repeat a subset on another date to distinguish a stable mode difference from ordinary live-search variation. Label the repeat as a new run, not an invisible extension.
Publish the manifest and limitations
Publish the query set where permissible, configuration fields, rubric, exclusions, raw or redacted captures, calculation code or formula, and observation dates. State that the study covers one account profile, locale, time window and product version. Keep the Kagi reproducibility guide beside the run and use the tool-score evaluation guide to expose any weights.
Do not call the AI-off condition “no algorithms” or “unpersonalized” unless those separate properties were controlled and documented. The switch concerns AI-based features, not every system involved in search.
Primary documentation
Source review: August 31, 2026. This protocol reports no comparison result until a declared sample is run.
Keep learning
Continue this topic
Next in this topic
Bing AI-Guided Image Search: What 8 Desktop Queries Showed
Earlier in this topic
Which Schema.org Types Are Actually Used Across the Web? July 2026 Data
Research
Ask a question or join the discussion