Page Weight vs Extractable Text: Results From a 25-Site Audit

We measured compressed response size, decoded HTML, server-extractable text and rendered text across 25 public sites. The results show why page weight alone is not a crawlability verdict.

Sonar routes one HTML page through server and rendered extraction gates, producing different text lengths that are clipped together for comparison.

Direct answer: in our 25-page, 25-domain audit, heavier HTML was not a reliable proxy for how much text a fixed browser run exposed. Decoded document size and rendered text had a moderate rank association (Spearman ρ = 0.452, n = 23), but almost no linear association (Pearson r = 0.004). A few very large product pages exposed far less visible text than smaller articles and documentation pages.

The practical lesson is narrower than “heavy pages are bad.” Save the response, define the server parser, render the page under a recorded environment, and retain failures. Download the 25-row audit dataset (CSV). It includes every frozen row, including the 403, zero-text and browser-failure observations.

The key result: bytes and visible text were not interchangeable

We selected five homepages, five articles, five product or service pages, five documentation pages and five directory or listing pages. Every row used a distinct registrable domain. The sample was frozen and hashed before measurement; a failed page could not be swapped for a cleaner result.

Initial-pass medians for the purposive sample
Page jobCompressed responseDecoded HTMLServer textRendered text
Homepage11.3 KiB52.8 KiB7,521 chars (n=4)4,655 chars (n=4)
Article48.4 KiB220.8 KiB20,992 chars17,729 chars
Product/service69.6 KiB453.0 KiB12,167 chars5,370 chars (n=4)
Documentation22.6 KiB179.1 KiB11,655 chars9,371 chars
Directory/listing9.4 KiB45.5 KiB2,400 chars2,803 chars

These are sample medians, not web-wide benchmarks. Product pages were the largest group by median decoded HTML, yet their median rendered text was below the article and documentation groups. Page job and representation mattered alongside size.

One page produced three different measurements

Compressed response bytes are the encoded body received by the HTTP collector. The Content-Encoding documentation explains why encoded and decoded sizes are distinct: gzip, Brotli or another content coding can shrink the transferred representation without changing the decoded HTML.

Server text came from the saved HTML after removing script, style, noscript, template and svg, then normalizing textContent. JavaScript from the measured page was not executed in this arm. MDN’s textContent reference describes descendant text extraction; our removals and whitespace rules are study-specific.

Rendered text came from document.body.innerText after Chromium reached DOMContentLoaded and waited five seconds. The HTML Standard defines innerText as text “as rendered.” We also saved the rendered DOM, a full-page screenshot and Chrome’s full accessibility tree. The accessibility endpoint is marked experimental in the Chrome DevTools Protocol, so the raw tree is an environment-specific artifact, not a stable cross-browser accessibility score.

All 25 frozen rows

The table reports the initial pass. “Repeat” describes the one verification attempt allowed by the preregistration. A missing rendered value remains missing; it is not imputed from another page or a later clean run.

Initial measurements for all 25 frozen URLs
Domain / IDPage jobHTTPCompressedDecoded HTMLServer text charsRendered text charsRepeat
searchengineanswer.com
H01
Homepage20038.7 KiB190.0 KiB8,4735,473Difference repeated
wikipedia.org
H02
Homepage20029.9 KiB117.5 KiB6,2011,992Difference repeated
python.org
H03
Homepage20011.2 KiB51.4 KiB6,5683,836Repeat timed out
w3.org
H04
Homepage4033.2 KiB5.7 KiB;;403 repeated
eff.org
H05
Homepage20011.3 KiB52.8 KiB8,5715,832Difference repeated
arxiv.org
A01
Article20042.6 KiB42.6 KiB4,8023,152Difference repeated
plato.stanford.edu
A02
Article200220.8 KiB220.8 KiB186,053186,186Not required
ourworldindata.org
A03
Article20017.8 KiB100.9 KiB7,6408,284Not required
redhat.com
A04
Article20048.4 KiB403.5 KiB20,99217,729Not required
ibm.com
A05
Article20051.6 KiB247.7 KiB28,49329,739Not required
slack.com
P01
Product/service20044.1 KiB212.7 KiB12,1674,404Difference repeated
anthropic.com
P02
Product/service200460.0 KiB1384.8 KiB49,3755,495Difference repeated
atlassian.com
P03
Product/service200142.8 KiB894.0 KiB05,245Zero/difference repeated
apple.com
P04
Product/service20069.6 KiB453.0 KiB42,08236,667Not required
adobe.com
P05
Product/service2005.3 KiB25.8 KiB1,457;Render error repeated
docs.github.com
D01
Documentation20025.8 KiB230.4 KiB11,3309,371Not required
developer.mozilla.org
D02
Documentation20020.3 KiB182.0 KiB11,6557,785Difference repeated
cloud.google.com
D03
Documentation20022.6 KiB112.0 KiB12,07710,608Not required
docs.aws.amazon.com
D04
Documentation20015.9 KiB60.8 KiB34,87936,587Not required
docs.docker.com
D05
Documentation20057.5 KiB179.1 KiB1,8151,495Not required
data.gov
L01
Directory/listing20033.5 KiB119.0 KiB2,4002,029Not required
nps.gov
L02
Directory/listing2006.8 KiB29.2 KiB1,0841,436Not required
usa.gov
L03
Directory/listing2009.4 KiB45.5 KiB3,3182,803Not required
cdc.gov
L04
Directory/listing2005.4 KiB27.1 KiB2,3664,328Difference repeated
un.org
L05
Directory/listing200104.1 KiB104.1 KiB17,36617,101Not required

The outliers explain why one ratio was not enough

  • Claude/Anthropic: 1,418,084 decoded HTML bytes, 49,375 server-text characters and 5,495 rendered-text characters. The counts repeated.
  • Jira: 915,469 decoded HTML bytes, zero server text under the registered parser, and 5,245 rendered characters initially (5,235 on repeat).
  • Apple MacBook Air: 463,896 decoded HTML bytes and 36,667 rendered characters. A large product document did not always produce little rendered text.
  • Stanford Encyclopedia of Philosophy: 226,148 decoded HTML bytes and about 186,000 characters in both extraction arms.
  • AWS S3 documentation: 62,278 decoded HTML bytes and 36,587 rendered characters.

These rows pulled in different directions. The rank relationship between decoded HTML and rendered text was positive, but the linear relationship collapsed around the extreme combinations. Reporting only a correlation;or only a “bytes per text character” ratio;would hide the page jobs, templates and failure modes that produced it.

Failures were part of the result

Rows that did not produce a complete two-arm observation
PageInitial observationVerificationReporting decision
W3C homepageHTTP 403 on measured GETHTTP 403 againKeep the frozen row; do not render or substitute
Adobe Creative CloudSaved response returned 200; Chromium navigation failed with ERR_HTTP2_PROTOCOL_ERRORSame browser errorReport server text; rendered value remains missing
Python.orgRendered 3,836 charactersNavigation timed outKeep both attempts as run instability
Jira200 response; zero server text; 5,245 rendered charactersZero server text; 5,235 rendered charactersKeep the zero; describe the representation gap

A failed collector run does not prove that a search crawler failed. It proves that this declared collector, from this environment, did not complete that layer. The distinction protects the result from becoming a crawler myth.

What the correlations mean;and do not mean

Exploratory associations inside this sample
PairnPearson rSpearman ρ
Decoded HTML vs server text240.1910.534
Decoded HTML vs rendered text230.0040.452
Compressed response vs server text240.5180.485
Server text vs rendered text230.9710.840

The server and rendered text counts moved together overall, but several large gaps repeated. That supports using the cheaper server extraction for triage only when a rendered check and exception rule remain available. It does not support declaring the server parser equivalent to a browser, Googlebot or an AI crawler.

How to reproduce the audit

  1. Freeze the URL and page-job list before the first measured request. Hash the file and keep failed URLs in place.
  2. Record user agent, region, time zone, browser/build, viewport, locale, wait rule, timeout and parser revision.
  3. Save the encoded response body, decoded HTML, headers without cookies, redirects, status and hashes.
  4. Parse the saved HTML without executing page JavaScript. Publish the removed elements and whitespace rule.
  5. Render the same URL in a clean browser context and save innerText, DOM, screenshot, accessibility tree, network events and errors.
  6. Predeclare when one repeat is allowed. Preserve both attempts even when the second looks cleaner.
  7. Review outliers before interpreting a ratio or correlation. Check templates, visibility, runtime state, redirects and collector failures.

Our fixed environment used Node 26.0.0, Playwright 1.60.0 and Chrome for Testing 148.0.7778.96 on macOS arm64. The browser waited five seconds after DOMContentLoaded; service workers were blocked. Full-page screenshots used Playwright’s documented full-page screenshot option.

Limits

  • The sample is purposive, small, English-oriented and collected from one host in one time window.
  • Each page represents one domain; no result estimates the distribution of all pages on that site or across the web.
  • Content negotiation, geolocation, consent interfaces, experiments, server load, bot handling and page updates can change a rerun.
  • The main, article, [role=main], body fallback is a study rule, not a universal main-content extractor.
  • Character counts are descriptive. They do not measure usefulness, originality, correctness, accessibility or business value.
  • The study did not use Googlebot, Bingbot or an AI crawler and did not measure indexing, ranking, retrieval, citation or traffic.

A better page-weight decision

Use page weight to find pages worth inspecting, not to announce a crawl failure. A useful decision record keeps five columns together: response bytes, decoded HTML bytes, server text, rendered text and collection status. Add a repeat only under the declared failure or difference rule.

If a page is large and both text arms are strong, optimize performance without inventing an extraction problem. If server text is empty but rendered text exists, inspect the source and runtime dependency. If rendering fails, reproduce the browser path before naming a crawler. If both arms contain little decisive content, the editorial or template question is more important than the byte count.

Result in one sentence

Page weight did not act as a substitute for extractable text

Across the frozen 25-site sample, compressed bytes, decoded HTML, server-extractable text and rendered text described different properties of a page. A large response could still expose strong text, while a smaller response could depend on rendering or fail for another reason.

  • Use transfer size for performance and delivery questions.
  • Use server text and rendered text as separate extraction arms.
  • Treat timeouts and browser failures as results that require diagnosis, not as zero-text pages.
Table comparing compressed response size, decoded HTML, server text and rendered text across page types
Original result table from the completed 25-site page-weight audit.

My takeaway: My practical conclusion is to stop using one byte-to-text ratio as a crawlability verdict. Inspect the layer that failed before recommending a fix.

Sources, method and data

Original data: 25 initial rows and 12 preregistered verification attempts collected August 31, 2026. The frozen sample, run manifest, row metadata, raw responses, rendered DOMs, accessibility trees, screenshots, network/error ledgers, hashes and analysis outputs are retained in the Search Engine Answer editorial evidence package.

Primary method sources: MDN documentation for Content-Encoding and textContent, the WHATWG HTML Standard for innerText, Chrome DevTools Protocol documentation for the accessibility tree, and Playwright documentation for browser page and screenshot capture; checked August 31, 2026.

Release note: this page was originally published as a testing-pending protocol. The findings above replace that placeholder only after the complete frozen sample and required repeats were preserved. The results apply to the declared pages, collector, environment and date.

Keep learning

Continue this topic

Community discussion

Discuss: Page Weight vs Extractable Text: Results From a 25-Site Audit

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.