// guide

What can you cut from a web page before your LLM stops understanding it? Six formats measured in Python

We turned 10 web pages into six LLM context formats, from raw HTML and Markdown to accessibility trees and screenshots, and measured token cost, read accuracy, and act accuracy. Fewer tokens rarely meant better answers.

Stagehand team10 min read
PythonLLM contextBrowser automation
A web page wireframe split into six encodings of decreasing size, with the highlighted ones feeding a model's context window
++

LLMs' data cutoff makes feeding them with web data a must-have. While hosted LLM APIs offer native web-search and web-fetch tools, they offer limited control and interactivity, leading us to resort to controlling a browser.

The issue is that LLMs have a limited context window that can easily be bloated with tool results. The real questions become: how to feed a web page to your LLM? Is raw HTML too much? Is extracted visible text as markdown not enough?

Turns out the answer is not binary: “it depends” (spoiler: accuracy doesn't track token count). In this article, we'll use Python to benchmark 6 different approaches to “represent” a web page to an LLM. So if you like numbers and code, welcome.

Six approaches to converting a page to LLM context#

The simplest way to read a page is to simply get the full HTML of a page and feed it to the LLM using page.content(). This comes with an issue, as over the years, webpages became heavier and heavier, filled with “ghost” <div> used by front-end framework to create rich user experiences but meaningless for semantics.

The logical next step is to only extract the text visible on screen and convert it to markdown. However, only extracting the visible text removes all their semantics (ex: UI action).

This is when more advanced approaches come into play, all looking to prune the DOM/HTML to only keep its meaningful semantic parts:

  • Accessibility tree, describing the purpose each meaningful HTML element using ARIA attributes is a common approach used by popular frameworks like Playwright.
  • Another approach used a Indexed interactive DOM, which filters interactive DOM elements and sort them by importance, an approach developed by the Browser Use library.
  • Stagehand uses a combination of the 2 former, producing a pruned DOM merged with the accessibility tree.

Finally, frontier models power their computer use capabilities by using page screenshots, feeding the LLM with images of the web page.

All the above 6 formats gives us the following list to benchmark:

FormatWhat it isHow to produce it in PythonWhat the model gets
Raw HTMLThe full page markup, as renderedpage.content()Every tag, attribute, script, and style
MarkdownThe page converted to Markdownmarkdownify on page.content()Headings, text, links, and tables, without markup
Accessibility treeThe browser's accessibility view of the pagepage.locator("body").aria_snapshot()Elements by role, name, and state (button, link, textbox)
Indexed interactive DOMA filtered DOM with numbered interactive elements, as used by Browser UseBrowser Use's DOM serializerClickable and fillable elements with indices to target them by
Pruned DOM + accessibility treeA pruned DOM merged with the accessibility tree, as used by StagehandStagehand's page snapshotPage structure and interactive elements with references to target them by
ScreenshotAn image of the visible viewportpage.screenshot()Pixels of what a user sees above the fold

Let's now look at the methodology used to evaluated these 6 web page to LLM context formats.

Benchmarking the 6 web pages representation formats#

Our 6 web pages representation formats are benchmarked against 10 web pages across 10 websites (books.toscrape.com, saucedemo.com, and local ones) representing common web tasks with different level of difficulties (forms, dashboards, an embedded iframe with a closed shadow root).

Here's the details of the 3 online web sites used during the benchmarks:

PageSourceWhy it's included
books_listingbooks.toscrape.comA real e-commerce grid: many repeated product cards, prices, links, and ratings. Tests how formats scale with repetition on a typical listing page.
books_productbooks.toscrape.comA real product detail page. Tests reading specific facts (price, stock, review count) from a dense page, including information stored only in CSS classes, like the star rating.
saucedemo_loginsaucedemo.comA real, minimal login form. Tests whether a format exposes input fields that rely on placeholders rather than visible labels.

The remaining 7 local web pages were built using the following fixtures:

PageReal-world pattern it modelsWhat it tests
articleBlog posts, news, documentation proseThe baseline: mostly text, few interactive elements. The page type where Markdown should do well.
docsDocumentation sites with navigation, version selectors, and collapsible sectionsNavigation-heavy pages, <select> dropdowns (version picker), <details>/<summary> elements
formSign-up, checkout, and settings formsForm state: selected options, pre-filled values, inputs without visible labels. The things a user sees but plain text formats drop.
dashboardSaaS admin panels and analytics UIsMixed content: filters, selected states, and an inline SVG chart with text inside it.
serpSearch results and any paginated listResult snippets, pagination links ("2", "3"...), and a <summary> element. Pagination is one of the most common agent actions.
embedEmbedded checkouts, payment forms, support widgets, cookie banners, web componentsSame-origin and cross-site iframes plus a closed shadow root. These are common on production sites (Stripe-style payment fields, chat widgets, design-system components) and invisible to standard formats.
tableData tables, reports, admin lists, spreadsheet-like views500 rows. Tests how each format scales on data-heavy pages, where token cost can grow faster than expected.

Our 6 representations are evaluated on 3 criteria: the size of the representation (in tokens), its read accuracy (how well the provided LLM representation describes the page) and its act accuracy (how well the model performed a given action with the given LLM representation).

The benchmark was run using one model (claude-sonnet-5) with 3 repeats and deterministic scoring, totaling for 1,440 calls (about $21.59).

Let's dig into the results, shall we?

How many tokens each representation outputs#

First, let's look at the median number of tokens generated by each representation:

PageRaw HTMLMarkdownA11y treeBrowser UseStagehandScreenshot
saucedemo_login1,0911071341924031,337
form3,1094751,1401,3932,4011,337
books_listing15,1184,9026,5274,1487,4161,337
table52,97730,08779,71973,27176,7021,337

Note: counts are from Anthropic's tokenizer; o200k counts run 17 to 34% lower.

For most scenarios, using markdown or an accessibility tree produces the smaller output representation of the medium-sized web page. The only exception is the 500 rows tables were markdown dominates (markdown has an efficient native way to represent tables).

On the other hand, the longer a page gets, the most efficient it is to just use a screenshot.

This gives us a great picture of how much tokens each representation cost but not how efficient they are at helping an LLM to understand a web page.

The twist: accuracy doesn't follow tokens#

That's the spoiler, reducing the number of tokens to represent a web page comes at a cost.

The table below shows how much reducing tokens reduces the overall LLM accuracy to understand and act on a web page:

FormatREADACTTokens per correct answer (excl. table)
Raw HTML92%79%5,094
Markdown80%70%1,992
Accessibility tree88%83%2,708
Indexed interactive DOM (Browser Use)86%97%1,965
Pruned DOM + accessibility tree (Stagehand)96%97%3,196
Screenshot72%73%1,992

Without surprise, providing raw HTML helps the LLM to understands a page, scoring a 92% on READ, however it makes it hard to act on it.

Interestingly, Stagehand's hybrid approach of pruning the DOM and merge it with the accessibility tree yields the higher accuracy scores on both READ and ACT, at the cost of a 1.5x higher token cost than a markdown representation.

Let's dig into the flaws of each representation:

  • Markdown loses state: It loses form state (selected options, pre-filled values) and inputs that have no visible label. It scores 6/24 on attribute-only tasks and 0/9 on the saucedemo login inputs.
  • Screenshots lose everything below the fold: 3/39 below the fold, 0/27 far below it, and 129/129 above the fold.
  • The accessibility tree loses iframes and closed shadow roots: Raw HTML, Markdown and the default aria_snapshot() can't see either iframe or the closed shadow root. They score 0/21 on those calls.
  • Indexed interactive DOM (Browser Use) loses single-character text and selected state: The serializer drops text nodes of one character or less, so the pagination links "2" to "9" appear as empty <a /> elements. And it lists <select> options without marking the selected one (3 READ tasks missed).
  • Pruned DOM + accessibility tree (Stagehand) costs more tokens: No coverage failure, but 1.58x to 2.45x Browser Use's tokens per page (median 1.74x), and more than the accessibility tree on 9 of 10 pages.

The answer is “it depends”#

There is no single-answer to our question “What can you cut from a web page before your LLM stops understanding it?”, but instead, a set of rule of thumbs depending on your use case.

If your Python program is primarily acting on web pages, don't rely on simple representation like raw HTML, markdown or screenshot. Instead use library like Stagehand that scores well on both reading and acting on the page (because to act on a web page, you need to first understand well).

If you're interested in quickly extracting data from short to medium pages, markdown, accessibility tree and screenshots are solid approaches.

If you're looking to accurately extract data, consider using raw HTML or libraries like Stagehand that offers the best read accuracy.

Conclusion#

The formats that saved the most tokens were not the ones that answered the most tasks correctly, because each one leaves out something different: form state, content below the fold, embedded widgets, or rows of a long table. Results on your own pages will depend on what your tasks need. The benchmark, fixtures, tasks, and raw logs are all in the public GitHub repository, and you can rerun it on your own pages and model with a single command.

++