Skip to Content
ReferenceMCP tools

MCP tools

The Lightpanda MCP server exposes these tools and resources, started with lightpanda mcp. See how to use MCP for practical documentation.

In the Arguments column, ? marks an optional argument, and selector / backendNodeId marks a pair where one of the two is required. Prefer selector for reproducibility; it also wins when you pass both. backendNodeId values come from a prior tree, links or findElement call.

Calling a tool

Every call goes through the MCP tools/call method, with name and arguments populated from the tables below.

{"jsonrpc":"2.0","id":2,"method":"tools/call", "params":{"name":"fill","arguments":{"selector":"#search","value":"lightpanda"}}}
{"result":{"content":[{"type":"text","text":"Filled element (selector: #search) with \"lightpanda\". Page url: https://example.com, HTTP 200 OK, title: Example Domain"}], "structuredContent":{"url":"https://example.com","httpStatus":200,"title":"Example Domain"},"isError":false}}

The result comes back as {"result":{"content":[{"type":"text","text":"..."}],"isError":false}}.

Page state. goto and the tools that act on the page (click, fill, scroll, hover, press, selectOption, setChecked) end their text with Page url: <url>, HTTP <status>, title: <title>; goto reports Navigated successfully. HTTP <status>. A 4xx or 5xx status means the page is an error page, not the content.

The same tools also return structuredContent with url, httpStatus and title, declared by each tool’s outputSchema in tools/list. Only url is always present; httpStatus is absent before any response arrives.

goto returns an error directly on a genuine navigation failure (e.g. navigation failed: CouldntResolveHost); a timeout reports "Navigation started (HTTP <status>) but the page did not finish loading before the timeout." (see Known behaviors).

These tools bring a page into the browser:

NameArgumentsDescription
gotourl, timeout?, waitUntil?Navigate to a specified URL and load the page in memory so it can be reused later for info extraction. Returns the HTTP status once waitUntil fires, or a timeout notice: a 4xx or 5xx means you got an error page, so check it before reading on. waitUntil accepts the same states as waitForState and defaults to load; prefer domcontentloaded plus a follow-up waitForSelector on pages whose late scripts (ads) hold load back, and avoid done on pages with constant background activity, since it can run to the timeout.
searchquery, timeout?Run a web search and return results as markdown: a numbered list of {title, url, snippet}. Search tries brave, tavily, exa, then keenable in order, each when its API key (BRAVE_API_KEY, TAVILY_API_KEY, EXA_API_KEY or KEENABLE_API_KEY) is set; keenable also works without a key through its public endpoint (rate-limited per client IP). Prefer this over goto-ing google.com/search directly (Google blocks the browser on User-Agent/TLS). The browser does not navigate; to open a result, use goto with its URL.

Reading the page

These tools read the loaded page without modifying it:

NameArgumentsDescription
markdownselector?, backendNodeId?, maxBytes?, strip?, url?, timeout?Render the page, or a subtree, as markdown. Scope with selector or backendNodeId to read just the relevant region: full-page markdown is the last resort. Use maxBytes to cap long pages. strip takes the same element groups as the html tool.
htmlselector?, backendNodeId?, maxBytes?, strip?, url?, timeout?Raw HTML for the document, or a single node’s outerHTML when scoped. Verbose; use only when you need attributes that markdown discards. Use maxBytes to cap long pages. strip is an object of element groups to omit: js (script, noscript, script preloads), css (style, stylesheet links), ui (css plus img, picture, video, audio, svg, canvas, iframe), invisible (elements set to display:none), shell (nav, aside, dialog, page-level header/footer and matching landmark roles; skipped when that would drop most of the text) and clutter (keep only the main content; includes shell and invisible, and falls back to shell when it finds too little). {"js":true,"css":true} keeps a page dump small.
screenshotpath?, selector?, backendNodeId?, fullPage?, strip?, url?, timeout?Render the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate browser rendering (no images, fonts or CSS colours). With path, writes the file at full size and returns its location; without it, returns the image inline where the client can display one, at most 1280px wide and 4096px tall. strip takes the same element groups as the html tool. Use it to see spatial layout; read content with markdown/tree.
treebackendNodeId?, maxDepth?, url?, timeout?Simplified semantic DOM tree: role, name, value, and backendNodeId per node. Pass backendNodeId to scope, maxDepth to limit depth.
linkslimit?, url?, timeout?Extract all links as text (visible anchor text, falling back to aria-label/title/image alt), href (resolved URL), and backendNodeId (pass to click/nodeDetails). One entry per href; hidden links are omitted. limit returns at most that many links, in document order.
nodeDetailsbackendNodeIdTag, role, name, interactivity, disabled, value, input type, placeholder, href, id, class, checked, and select options for a node, plus a ready-to-use CSS selector that resolves to the node (the first match, as click/fill resolve it). The canonical way to turn a tree backendNodeId into a CSS selector.
findElementrole?, name?Find interactive elements by role and/or accessible name. name matches case-insensitively as a substring, or as a JavaScript regex literal such as /sign (in|up)/ (unanchored; flags i, m, s, u accepted; case-insensitive even without i, prefix (?-i) to make it case-sensitive). Returns matching elements with their backend node IDs. Useful for locating specific elements without parsing the full semantic tree.
interactiveElementsurl?, timeout?List every visible interactive element as a JSON array: native controls, ARIA widgets, contenteditable regions, elements with event listeners, and focusable elements. Each entry has backendNodeId, tagName, role, name, type (why it counts as interactive), tabIndex, and when present listeners, disabled, id, class, href, inputType, value, elementName and placeholder. To locate one element by role or name, findElement is cheaper.
structuredDataurl?, timeout?Page metadata as JSON: jsonLd (each JSON-LD block as a string), openGraph, twitterCard, meta and links (key/value lists), plus alternate (hreflang variants) and linkHeaders (relations from the HTTP Link header) when present. Empty sections come back as empty arrays. Use it for publisher-declared facts such as product price, article author or canonical URL before scraping the visible text.
detectFormsurl?, timeout?List the forms on the page as JSON: each form’s backendNodeId, action, method and fields, where each field has backendNodeId, tagName, name, inputType, required, disabled, and when present value, placeholder and select options. It returns no CSS selectors; get one per field with nodeDetails so fill calls stay replayable.

Calling a read tool with url set is more efficient than a separate goto: it navigates and reads in one request.

Data extraction and scripting

These tools return structured results from the loaded page:

NameArgumentsDescription
extractschema, save?Extract structured data from the current page (navigate first) using a schema mapping output field names to CSS-selector specs.
evaluatescript, url?, timeout?, save?Evaluate JavaScript in the page context. A bare trailing expression yields its value; top-level await and return are supported. This is an escape hatch for page-side logic the dedicated tools can’t express: prefer extract for data and click/fill/etc. for actions. It runs in the page, so it cannot see the agent script’s variables or builtins: interpolate any value into the script string. Objects and arrays return as JSON, so no JSON.stringify is needed. If a url is provided, it navigates there first. The globalThis.lp object exposes a Session-scoped bridge store: values written via lp.foo = ... auto-sync at end of evaluate, surviving navigation; values previously set via /extract save= or /evaluate save= appear as lp.<name>.

save stores the result under a bridge key, available in later evaluate calls as lp.<name>.

extract schema

schema is a JSON object (passed as a string) mapping output field names to CSS-selector specs. It is NOT a JSON Schema: no "type"/"properties" wrappers; the keys ARE your output fields. Value shapes:

Value shapeResult
"<sel>"First match’s text (trimmed; null if no match)
["<sel>"]Every match’s text (string[])
{"selector":"<sel>","attr":"<name>"}First match’s attribute value (href/src resolved to absolute URLs)
[{"selector":"<sel>","attr":"<name>"}]Every match’s attribute (string[])
[{"selector":"<sel>","fields":{…}}]One object per match; field selectors resolve relative to that match and accept any shape above ("" = the match’s own text; nest arrays for per-item sub-lists)

Add "limit": N inside any array’s object spec to cap matches. Every extracted value is a string or null; parse numbers downstream. An empty array is a valid result, but if ALL top-level keys miss, the call errors: inspect the page (tree/markdown) and retry with corrected selectors. Finish data tasks with extract: it is the only read recorded as a replayable extract(...) script call; answers lifted from markdown text in chat are not.

Examples (schema → result):

SchemaResult
{"karma": "#karma"}{"karma":"42"}
{"items": [".story .title"]}{"items":["Title 1","Title 2"]}
{"top3": [{"selector":".story .title","limit":3}]}{"top3":["A","B","C"]}
{"links": [{"selector":"a.title","attr":"href"}]}{"links":["https://site/a","https://site/b"]}
{"stories": [{"selector":".athing","fields":{"title":".titleline","rank":".rank"}}]}{"stories":[{"title":"Foo","rank":"1"}]}

Interacting with the page

These tools dispatch real DOM events on the page:

NameArgumentsDescription
clickselector / backendNodeIdClick on an interactive element. Returns the current page URL and title after the click.
fillselector / backendNodeId, valueFill text into an input element. Returns the filled value and current page URL and title.
scrollselector? / backendNodeId?, x?, y?Scroll the window, or an element’s scroll container, to an absolute position; an omitted axis keeps its current offset. Target an element with selector (preferred) or backendNodeId; omit both to scroll the window. If that element isn’t itself a scroll container, its nearest scrollable ancestor scrolls instead, and the result reports whichever node actually moved. Page scripts receive a scroll event, so content that loads on scroll may appear: read the page again afterwards, with waitForState if it is still loading. Returns the resulting scroll position and the current page URL and title.
hoverselector / backendNodeIdHover over an element, triggering mouseover and mouseenter events. Useful for menus, tooltips, and hover states.
presskey, selector?, backendNodeId?Press a keyboard key, dispatching keydown and keyup events. Use key names like ‘Enter’, ‘Tab’, ‘Escape’, ‘ArrowDown’, ‘Backspace’, or single characters like ‘a’, ‘1’. Common shorthand is normalized: ‘enter’/‘return’ → ‘Enter’, ‘esc’ → ‘Escape’, ‘up’/‘down’/‘left’/‘right’ → ‘Arrow*’, ‘space’ → ’ ’. Pressing ‘Enter’ on a form input or submit button triggers implicit form submission. Targets the document if no element is given.
selectOptionselector / backendNodeId, valueSelect an option in a <select> dropdown element by its value. Dispatches input and change events.
setCheckedselector / backendNodeId, checkedCheck (true) or uncheck (false) a checkbox or radio button. Dispatches input, change, and click events.

Waiting

These tools block until the page reaches a condition:

NameArgumentsDescription
waitForSelectorselector, timeout?Wait for an element matching a CSS selector to appear in the page. Returns the backend node ID of the matched element.
waitForScriptscript, timeout?Wait until a JavaScript expression returns truthy, re-evaluating on each tick of the event loop. Use for synchronization beyond what CSS selectors can express: e.g. window.dataLoaded === true, document.readyState === 'complete', document.querySelectorAll('.row').length >= 5.
waitForStatestate, timeout?Wait for the page to reach a load state (load, domcontentloaded, networkalmostidle, networkidle, or done), with no navigation. After a goto, the page is returned at the fast load snapshot, so content rendered by post-load JS (XHR-loaded lists, feeds, search results) may still be missing. When a read looks incomplete: empty lists, spinners, skeletons; call this with ‘networkidle’ and re-read. Prefer ‘networkidle’; ‘done’ can be slow on sites with constant background activity (ads, polling).

State and debugging

These tools inspect browser state outside the DOM:

NameArgumentsDescription
getUrl—Get the URL of the page currently loaded in the browser. Useful to verify a navigation or detect a redirect.
getCookiesurl?, all?Get cookies stored in the browser. Defaults to the current page’s host; pass url to filter another host or all to dump every cookie. Useful for debugging authentication and session state.
getEnvname?With name: read an LP_* environment variable, for non-secret config only (base URLs, flags). Without name: list the LP_* names that are set, for safe credential discovery. For secrets, pass $LP_* placeholders in tool args; never request a credential by name.
consoleLogs—Get buffered console.log/warn/error messages from the current page, then clear the buffer.

Session

session_new, session_list and session_close require the HTTP transport (lightpanda mcp --port <PORT>); over stdio only the default session exists and these calls return an error. Over HTTP transport, every other tool operates on whichever session the Mcp-Session-Id header names, falling back to the always-present default session.

NameArgumentsDescription
savepath, scriptSave the session as a reusable PandaScript (.js).
session_newname?Create a new isolated browser session (its own page, cookies and memory) and return its id. Use it to give a separate agent its own browsing context, or to obtain an id to share. Pass that id back as the Mcp-Session-Id header to route calls to it.
session_list—List the active browser sessions with their id and current URL. The default session always exists.
session_closeidClose a browser session, freeing its page and memory. The default session cannot be closed.

Resources

Three read-only resources are available, read via resources/read. The two page resources need a loaded page; mcp://skill/pandascript doesn’t.

URIMIME typeDescription
mcp://page/htmltext/htmlThe serialized HTML DOM of the current page
mcp://page/markdowntext/markdownThe token-efficient markdown representation of the current page (identical output to the markdown tool)
mcp://skill/pandascripttext/markdownThe PandaScript skill documentation
{"jsonrpc":"2.0","id":2,"method":"resources/read", "params":{"uri":"mcp://page/markdown"}}

The markdown tool and the mcp://page/markdown resource return the same content. The difference is who initiates: tools are called by the agent during its workflow; resources are read by the host application (e.g. an IDE displaying page state in the background).