Page Weight vs Extractable Text: Results From a 25-Site Audit
We measured compressed response size, decoded HTML, server-extractable text and rendered text across 25 public sites. The results show why page weight alone is not a crawlability verdict.
Direct answer: in our 25-page, 25-domain audit, heavier HTML was not a reliable proxy for how much text a fixed browser run exposed. Decoded document size and rendered text had a moderate rank association (Spearman ρ = 0.452, n = 23), but almost no linear association (Pearson r = 0.004). A few very large product pages exposed far less visible text than smaller articles and documentation pages.
The practical lesson is narrower than “heavy pages are bad.” Save the response, define the server parser, render the page under a recorded environment, and retain failures. Download the 25-row audit dataset (CSV). It includes every frozen row, including the 403, zero-text and browser-failure observations.
The key result: bytes and visible text were not interchangeable
We selected five homepages, five articles, five product or service pages, five documentation pages and five directory or listing pages. Every row used a distinct registrable domain. The sample was frozen and hashed before measurement; a failed page could not be swapped for a cleaner result.
| Page job | Compressed response | Decoded HTML | Server text | Rendered text |
|---|---|---|---|---|
| Homepage | 11.3 KiB | 52.8 KiB | 7,521 chars (n=4) | 4,655 chars (n=4) |
| Article | 48.4 KiB | 220.8 KiB | 20,992 chars | 17,729 chars |
| Product/service | 69.6 KiB | 453.0 KiB | 12,167 chars | 5,370 chars (n=4) |
| Documentation | 22.6 KiB | 179.1 KiB | 11,655 chars | 9,371 chars |
| Directory/listing | 9.4 KiB | 45.5 KiB | 2,400 chars | 2,803 chars |
These are sample medians, not web-wide benchmarks. Product pages were the largest group by median decoded HTML, yet their median rendered text was below the article and documentation groups. Page job and representation mattered alongside size.
One page produced three different measurements
Compressed response bytes are the encoded body received by the HTTP collector. The Content-Encoding documentation explains why encoded and decoded sizes are distinct: gzip, Brotli or another content coding can shrink the transferred representation without changing the decoded HTML.
Server text came from the saved HTML after removing script, style, noscript, template and svg, then normalizing textContent. JavaScript from the measured page was not executed in this arm. MDN’s textContent reference describes descendant text extraction; our removals and whitespace rules are study-specific.
Rendered text came from document.body.innerText after Chromium reached DOMContentLoaded and waited five seconds. The HTML Standard defines innerText as text “as rendered.” We also saved the rendered DOM, a full-page screenshot and Chrome’s full accessibility tree. The accessibility endpoint is marked experimental in the Chrome DevTools Protocol, so the raw tree is an environment-specific artifact, not a stable cross-browser accessibility score.
All 25 frozen rows
The table reports the initial pass. “Repeat” describes the one verification attempt allowed by the preregistration. A missing rendered value remains missing; it is not imputed from another page or a later clean run.
| Domain / ID | Page job | HTTP | Compressed | Decoded HTML | Server text chars | Rendered text chars | Repeat |
|---|---|---|---|---|---|---|---|
| searchengineanswer.com H01 | Homepage | 200 | 38.7 KiB | 190.0 KiB | 8,473 | 5,473 | Difference repeated |
| wikipedia.org H02 | Homepage | 200 | 29.9 KiB | 117.5 KiB | 6,201 | 1,992 | Difference repeated |
| python.org H03 | Homepage | 200 | 11.2 KiB | 51.4 KiB | 6,568 | 3,836 | Repeat timed out |
| w3.org H04 | Homepage | 403 | 3.2 KiB | 5.7 KiB | ; | ; | 403 repeated |
| eff.org H05 | Homepage | 200 | 11.3 KiB | 52.8 KiB | 8,571 | 5,832 | Difference repeated |
| arxiv.org A01 | Article | 200 | 42.6 KiB | 42.6 KiB | 4,802 | 3,152 | Difference repeated |
| plato.stanford.edu A02 | Article | 200 | 220.8 KiB | 220.8 KiB | 186,053 | 186,186 | Not required |
| ourworldindata.org A03 | Article | 200 | 17.8 KiB | 100.9 KiB | 7,640 | 8,284 | Not required |
| redhat.com A04 | Article | 200 | 48.4 KiB | 403.5 KiB | 20,992 | 17,729 | Not required |
| ibm.com A05 | Article | 200 | 51.6 KiB | 247.7 KiB | 28,493 | 29,739 | Not required |
| slack.com P01 | Product/service | 200 | 44.1 KiB | 212.7 KiB | 12,167 | 4,404 | Difference repeated |
| anthropic.com P02 | Product/service | 200 | 460.0 KiB | 1384.8 KiB | 49,375 | 5,495 | Difference repeated |
| atlassian.com P03 | Product/service | 200 | 142.8 KiB | 894.0 KiB | 0 | 5,245 | Zero/difference repeated |
| apple.com P04 | Product/service | 200 | 69.6 KiB | 453.0 KiB | 42,082 | 36,667 | Not required |
| adobe.com P05 | Product/service | 200 | 5.3 KiB | 25.8 KiB | 1,457 | ; | Render error repeated |
| docs.github.com D01 | Documentation | 200 | 25.8 KiB | 230.4 KiB | 11,330 | 9,371 | Not required |
| developer.mozilla.org D02 | Documentation | 200 | 20.3 KiB | 182.0 KiB | 11,655 | 7,785 | Difference repeated |
| cloud.google.com D03 | Documentation | 200 | 22.6 KiB | 112.0 KiB | 12,077 | 10,608 | Not required |
| docs.aws.amazon.com D04 | Documentation | 200 | 15.9 KiB | 60.8 KiB | 34,879 | 36,587 | Not required |
| docs.docker.com D05 | Documentation | 200 | 57.5 KiB | 179.1 KiB | 1,815 | 1,495 | Not required |
| data.gov L01 | Directory/listing | 200 | 33.5 KiB | 119.0 KiB | 2,400 | 2,029 | Not required |
| nps.gov L02 | Directory/listing | 200 | 6.8 KiB | 29.2 KiB | 1,084 | 1,436 | Not required |
| usa.gov L03 | Directory/listing | 200 | 9.4 KiB | 45.5 KiB | 3,318 | 2,803 | Not required |
| cdc.gov L04 | Directory/listing | 200 | 5.4 KiB | 27.1 KiB | 2,366 | 4,328 | Difference repeated |
| un.org L05 | Directory/listing | 200 | 104.1 KiB | 104.1 KiB | 17,366 | 17,101 | Not required |
The outliers explain why one ratio was not enough
- Claude/Anthropic: 1,418,084 decoded HTML bytes, 49,375 server-text characters and 5,495 rendered-text characters. The counts repeated.
- Jira: 915,469 decoded HTML bytes, zero server text under the registered parser, and 5,245 rendered characters initially (5,235 on repeat).
- Apple MacBook Air: 463,896 decoded HTML bytes and 36,667 rendered characters. A large product document did not always produce little rendered text.
- Stanford Encyclopedia of Philosophy: 226,148 decoded HTML bytes and about 186,000 characters in both extraction arms.
- AWS S3 documentation: 62,278 decoded HTML bytes and 36,587 rendered characters.
These rows pulled in different directions. The rank relationship between decoded HTML and rendered text was positive, but the linear relationship collapsed around the extreme combinations. Reporting only a correlation;or only a “bytes per text character” ratio;would hide the page jobs, templates and failure modes that produced it.
Failures were part of the result
| Page | Initial observation | Verification | Reporting decision |
|---|---|---|---|
| W3C homepage | HTTP 403 on measured GET | HTTP 403 again | Keep the frozen row; do not render or substitute |
| Adobe Creative Cloud | Saved response returned 200; Chromium navigation failed with ERR_HTTP2_PROTOCOL_ERROR | Same browser error | Report server text; rendered value remains missing |
| Python.org | Rendered 3,836 characters | Navigation timed out | Keep both attempts as run instability |
| Jira | 200 response; zero server text; 5,245 rendered characters | Zero server text; 5,235 rendered characters | Keep the zero; describe the representation gap |
A failed collector run does not prove that a search crawler failed. It proves that this declared collector, from this environment, did not complete that layer. The distinction protects the result from becoming a crawler myth.
What the correlations mean;and do not mean
| Pair | n | Pearson r | Spearman ρ |
|---|---|---|---|
| Decoded HTML vs server text | 24 | 0.191 | 0.534 |
| Decoded HTML vs rendered text | 23 | 0.004 | 0.452 |
| Compressed response vs server text | 24 | 0.518 | 0.485 |
| Server text vs rendered text | 23 | 0.971 | 0.840 |
The server and rendered text counts moved together overall, but several large gaps repeated. That supports using the cheaper server extraction for triage only when a rendered check and exception rule remain available. It does not support declaring the server parser equivalent to a browser, Googlebot or an AI crawler.
How to reproduce the audit
- Freeze the URL and page-job list before the first measured request. Hash the file and keep failed URLs in place.
- Record user agent, region, time zone, browser/build, viewport, locale, wait rule, timeout and parser revision.
- Save the encoded response body, decoded HTML, headers without cookies, redirects, status and hashes.
- Parse the saved HTML without executing page JavaScript. Publish the removed elements and whitespace rule.
- Render the same URL in a clean browser context and save
innerText, DOM, screenshot, accessibility tree, network events and errors. - Predeclare when one repeat is allowed. Preserve both attempts even when the second looks cleaner.
- Review outliers before interpreting a ratio or correlation. Check templates, visibility, runtime state, redirects and collector failures.
Our fixed environment used Node 26.0.0, Playwright 1.60.0 and Chrome for Testing 148.0.7778.96 on macOS arm64. The browser waited five seconds after DOMContentLoaded; service workers were blocked. Full-page screenshots used Playwright’s documented full-page screenshot option.
Limits
- The sample is purposive, small, English-oriented and collected from one host in one time window.
- Each page represents one domain; no result estimates the distribution of all pages on that site or across the web.
- Content negotiation, geolocation, consent interfaces, experiments, server load, bot handling and page updates can change a rerun.
- The
main,article,[role=main],bodyfallback is a study rule, not a universal main-content extractor. - Character counts are descriptive. They do not measure usefulness, originality, correctness, accessibility or business value.
- The study did not use Googlebot, Bingbot or an AI crawler and did not measure indexing, ranking, retrieval, citation or traffic.
A better page-weight decision
Use page weight to find pages worth inspecting, not to announce a crawl failure. A useful decision record keeps five columns together: response bytes, decoded HTML bytes, server text, rendered text and collection status. Add a repeat only under the declared failure or difference rule.
If a page is large and both text arms are strong, optimize performance without inventing an extraction problem. If server text is empty but rendered text exists, inspect the source and runtime dependency. If rendering fails, reproduce the browser path before naming a crawler. If both arms contain little decisive content, the editorial or template question is more important than the byte count.
Result in one sentence
Page weight did not act as a substitute for extractable text
Across the frozen 25-site sample, compressed bytes, decoded HTML, server-extractable text and rendered text described different properties of a page. A large response could still expose strong text, while a smaller response could depend on rendering or fail for another reason.
- Use transfer size for performance and delivery questions.
- Use server text and rendered text as separate extraction arms.
- Treat timeouts and browser failures as results that require diagnosis, not as zero-text pages.

My takeaway: My practical conclusion is to stop using one byte-to-text ratio as a crawlability verdict. Inspect the layer that failed before recommending a fix.
Sources, method and data
Original data: 25 initial rows and 12 preregistered verification attempts collected August 31, 2026. The frozen sample, run manifest, row metadata, raw responses, rendered DOMs, accessibility trees, screenshots, network/error ledgers, hashes and analysis outputs are retained in the Search Engine Answer editorial evidence package.
Primary method sources: MDN documentation for Content-Encoding and textContent, the WHATWG HTML Standard for innerText, Chrome DevTools Protocol documentation for the accessibility tree, and Playwright documentation for browser page and screenshot capture; checked August 31, 2026.
Release note: this page was originally published as a testing-pending protocol. The findings above replace that placeholder only after the complete frozen sample and required repeats were preserved. The results apply to the declared pages, collector, environment and date.
Keep learning
Continue this topic
Next in this topic
ChatGPT Referral Growth: Separate AEO Lift from Platform Tailwind
Earlier in this topic
Accessibility Tree, HTML, or Screenshot? What 10 Fixed Tasks Preserved
Research
Ask a question or join the discussion