You.com Search Agents Can Choose Cached or Live Page Content
You.com's extraction_source setting exposes a practical freshness decision: use cached content for speed, live fetch for recency, or blend the two. The right choice depends on the claim's change rate.
Direct answer: You.com’s content extraction can use cached content, always fetch a live page, or blend both. The setting exposes a choice that search-agent builders often leave implicit: how much latency and cost should be exchanged for fresher evidence?
The answer should be made by claim, not by domain. A product specification may remain stable for months while its price, stock or service status changes within minutes.
The three extraction modes
You.com’s page-content retrieval guide documents an extraction_source field with three values. blend is the default and uses cached content where available, then live-crawls other pages. cache uses cached content only. fetch always requests live page content.
Cache is the fastest mode but content can be absent or older. Fetch is designed for freshness but is the slowest. Blend is a practical default when the result set contains both cached and uncached pages.
Freshness is a field-level requirement
| Claim | Typical change rate | Starting mode |
|---|---|---|
| Historical founding date | Rarely | Cache |
| Software feature documentation | Weeks or days | Blend |
| Product price or stock | Hours or minutes | Fetch |
| Service or incident status | Minutes | Fetch |
These are starting rules, not guarantees. The source itself may publish slowly or omit a timestamp, so a live fetch can still return stale information.
Live extraction has a distinct cost path
You.com’s documentation states that only pages crawled live incur the extraction charge. As checked on September 11, 2026, the documented rate is $1 per 1,000 live-crawled pages, in addition to the base search call. Cached extraction does not add that live-page fee. Pricing can change, so confirm the current You.com pricing before forecasting production cost.
Measure cost per usable current field. A live page that lacks the requested price or timestamp is billed activity, not a successful freshness result.
Run matched requests across all three modes
- Select URLs with known update patterns and machine-readable timestamps where possible.
- Record the source value and timestamp immediately before the test.
- Request the same URL and field with cache, blend and fetch.
- Measure latency, content presence and agreement with the known value.
- Repeat after a controlled source-page update.
- Record which requests caused a live crawl and their cost.
- Separate extraction failure from answer-generation failure.
A controlled update creates a clear freshness boundary. Without it, two modes can agree simply because the source did not change.
A two-stage policy limits unnecessary live fetches
Start with cache or blend for discovery and stable evidence. Escalate to fetch when a required field is missing, the cached timestamp exceeds the field’s freshness budget, or the claim belongs to a high-volatility class.
The policy should be deterministic. “Fetch when freshness matters” is too vague for auditing. “Fetch price when the content timestamp is older than six hours” can be tested and revised.
Do not collapse missing, stale and contradictory
- Missing: no extracted content or required field was returned.
- Stale: the field conflicts with a newer timestamped source value.
- Contradictory: two current sources report different values.
- Unverifiable: the result has no timestamp or source evidence strong enough to judge.
Each state needs a different fallback. A live refetch may solve stale content but cannot resolve two authoritative sources that disagree.
Publishers can test how quickly corrections propagate
When a public page changes a material fact, save the old and new value, update timestamp and exact URL. Query the extraction modes on a declared schedule until they return the new value. This produces a propagation observation for that URL and field, not a universal cache-expiry guarantee.
Pair the result with our crawler log verification guide if you control the server, and use the stable sitemap URL workflow to avoid confusing content freshness with changing discovery URLs.
Download the freshness test sheet
Download the extraction freshness CSV. It records the source value and timestamp, extraction mode, returned value, latency, content presence, agreement and live-page billing.
The EXAMPLE-REMOVE row is a schema example, not a measured You.com result.
The practical decision rule
Use cache for stable facts and high-volume discovery, fetch for time-sensitive fields where an older value would change the answer, and blend when the query contains both. Then measure the result rather than assuming the label describes the content’s actual age.
The best system is not the one that fetches most often. It is the one that meets a stated freshness requirement with a visible failure path and acceptable cost.
Primary documentation
Keep learning
Continue this topic
Next in this topic
Two Google Ads Changes Alter Search Campaign Controls
Earlier in this topic
ChatGPT Data Agent Makes Metric Definitions Part of SEO Reporting
Tools & Workflows
Ask a question or join the discussion