How to give your AI app live web data
A model only knows what it was trained on and what you put in front of it. To answer questions about today - a price, a policy, a listing, a changelog - it needs to read the live page. This guide shows how to connect an AI agent, RAG pipeline or LLM app to public web pages: how to build a fetch tool, how to clean what comes back, and how to keep your sources fresh.
Why models need live data
Training data has a cutoff, and a retrieval index is only as fresh as its last crawl. When someone asks about something that changed since then, a model without web access has two options: say it doesn't know, or guess. Live access lets it read the page instead.
The catch is that the pages you want are often the hardest to read automatically. Content built with JavaScript arrives as an empty shell in a plain HTTP request, and many sites turn away automated traffic. That is the layer UnblockingAPI handles: it loads the page in a real browser and returns what a visitor would see.
Three ways to connect
- — An assistant you already use. Add the MCP server to Claude, Cursor, VS Code or Zed and the assistant can fetch pages itself. No code.
- — A tool in your own agent. Define a fetch tool, let the model call it, and return the cleaned page. Covered in this guide.
- — A refresh job for RAG. Re-fetch the pages behind your index on a schedule and update only what changed.
First request
Send the page URL with rendering on. The response field of the JSON reply holds the HTML. See the API documentation for every parameter.
curl \
-H "X-Api-Key: YOUR_API_KEY" \
"https://api.unblockingapi.com/unblock?url=https://example.com/pricing&render=true"If you're not sure why a page comes back empty without rendering, read how to scrape JavaScript-rendered pages.
Build a fetch tool
A tool is a name, a description the model reads, and a schema for its input. This example uses the JSON-schema style most model APIs accept - adapt the wrapper to the SDK you use. The handler is plain code that runs when the model asks for the tool.
export const fetchPageTool = {
name: "fetch_page",
description:
"Fetch a public web page and return its readable text. Use when you need current information from a specific URL.",
input_schema: {
type: "object",
properties: {
url: { type: "string", description: "The full https:// URL to read" },
},
required: ["url"],
},
};
export async function fetchPage({ url }) {
const res = await fetch(
`https://api.unblockingapi.com/unblock?url=${encodeURIComponent(url)}&render=true`,
{ headers: { "X-Api-Key": process.env.UNBLOCKINGAPI_KEY } }
);
const data = await res.json();
return htmlToText(data.response);
}Your agent loop passes fetchPageTool to the model, and when the model replies with a request to call fetch_page, you run fetchPage and send the result back as the tool's output.
Clean the page for the model
Rendered HTML is mostly scripts, styles and navigation. Sending all of it wastes tokens and buries the content. A small cleaner using Cheerio keeps the visible text and caps the length:
import * as cheerio from "cheerio";
export function htmlToText(html, maxChars = 12000) {
const $ = cheerio.load(html);
$("script, style, noscript, svg, nav, footer, iframe").remove();
const text = ($("main").text() || $("body").text())
.replace(/\s+/g, " ")
.trim();
return text.length > maxChars ? text.slice(0, maxChars) + " …[truncated]" : text;
}Keep the source URL next to the text you return so the model - and your users - can cite where an answer came from.
Structured output
When the model only needs a few fields - a price, a title, a list of items - pass those instead of the whole page. You can build the extraction without writing selectors: open the page in the visual editor, tap the fields you want and save a template that returns clean JSON on every call. Ready-made options are in the templates library.
Keep RAG sources fresh
A retrieval index drifts out of date as the pages behind it change. A refresh job fetches each source on a schedule, hashes the cleaned text and only re-embeds pages whose hash changed:
import { createHash } from "node:crypto";
export async function refreshSource(source, store) {
const text = await fetchPage({ url: source.url });
const hash = createHash("sha256").update(text).digest("hex");
if (hash === source.lastHash) return { url: source.url, changed: false };
await store.reEmbed(source.url, text); // your vector store
await store.save(source.url, { hash, fetchedAt: new Date().toISOString() });
return { url: source.url, changed: true };
}Each fetch is a normal request, so you control the cost by choosing the schedule. For a longer look at monitoring pages over time, see the competitor research use case.
Handle failures honestly
Some fetches will fail or return a page with nothing useful on it. The important thing is what your tool tells the model. If it returns an empty string, the model may carry on as if it had read something. Return an explicit error instead:
const text = await fetchPage({ url });
if (!text || text.length < 200) {
return "ERROR: the page returned no readable content. Do not guess - tell the user you could not read it.";
}
return text;Assistants over MCP
If you'd rather not write a tool at all, the open-source MCP server gives an assistant the same ability out of the box. Setup guides: Claude, Cursor, VS Code and Zed. The wider picture is on the AI agents and LLM apps page.
Responsible use
Frequently asked questions
Give the model a tool that fetches a URL and returns the page content. When the model needs current information it calls the tool, your code fetches the page through UnblockingAPI, and the cleaned content goes back into the conversation. Assistants like Claude and Cursor can do this over MCP without any code.
Many sites build their content with JavaScript, so a plain HTTP request returns an empty shell, and some refuse automated requests. Fetching through a real browser returns the rendered page instead.
Usually not. Raw HTML is mostly markup and scripts, which costs tokens and distracts the model. Strip scripts and styles and pass the visible text, or extract only the fields you need as JSON.
It gives the model the actual page to answer from, which reduces answers invented from stale or missing information. It is not a guarantee, so keep the source URL with the content and let users check it.
As often as the underlying content changes. Fast-moving pages like prices or news may need daily or hourly refreshes; documentation may only need weekly. Compare a content hash and re-embed only the pages that actually changed.
No. UnblockingAPI is built for fetching specific public pages for live, per-request use. Use it within our Acceptable Use Policy and the terms of the sites you fetch.
Give your agent a page to read.
Test the fetching layer before building your parser. Start with 500 free credits and no credit card. See pricing for plan details.
One successful request is one credit. No multipliers for rendering, proxies or retries.
Related guidesBrowse all guides
How to scrape public property listing pages
Fetch JavaScript-rendered property listing pages with country-level routing and parse the returned HTML.
How to scrape e-commerce product pages and pricing data
Fetch public product pages through UnblockingAPI, parse pricing and product details from the returned HTML.
How to scrape JavaScript-rendered pages
Use real browser rendering to fetch pages that load content dynamically through client-side JavaScript.
Website authority checker: track authority signals at scale
What domain authority and authority score measure, and how to collect and track the public signals behind them - plus an analytics tag checker.
Last updated October 2026