Developer guide

How to give your AI app live web data

A model only knows what it was trained on and what you put in front of it. To answer questions about today - a price, a policy, a listing, a changelog - it needs to read the live page. This guide shows how to connect an AI agent, RAG pipeline or LLM app to public web pages: how to build a fetch tool, how to clean what comes back, and how to keep your sources fresh.

Practical guide 12 min read

Why models need live data

Training data has a cutoff, and a retrieval index is only as fresh as its last crawl. When someone asks about something that changed since then, a model without web access has two options: say it doesn't know, or guess. Live access lets it read the page instead.

The catch is that the pages you want are often the hardest to read automatically. Content built with JavaScript arrives as an empty shell in a plain HTTP request, and many sites turn away automated traffic. That is the layer UnblockingAPI handles: it loads the page in a real browser and returns what a visitor would see.

Three ways to connect

  • — An assistant you already use. Add the MCP server to Claude, Cursor, VS Code or Zed and the assistant can fetch pages itself. No code.
  • — A tool in your own agent. Define a fetch tool, let the model call it, and return the cleaned page. Covered in this guide.
  • — A refresh job for RAG. Re-fetch the pages behind your index on a schedule and update only what changed.

First request

Send the page URL with rendering on. The response field of the JSON reply holds the HTML. See the API documentation for every parameter.

curl \
  -H "X-Api-Key: YOUR_API_KEY" \
  "https://api.unblockingapi.com/unblock?url=https://example.com/pricing&render=true"

If you're not sure why a page comes back empty without rendering, read how to scrape JavaScript-rendered pages.

Build a fetch tool

A tool is a name, a description the model reads, and a schema for its input. This example uses the JSON-schema style most model APIs accept - adapt the wrapper to the SDK you use. The handler is plain code that runs when the model asks for the tool.

export const fetchPageTool = {
  name: "fetch_page",
  description:
    "Fetch a public web page and return its readable text. Use when you need current information from a specific URL.",
  input_schema: {
    type: "object",
    properties: {
      url: { type: "string", description: "The full https:// URL to read" },
    },
    required: ["url"],
  },
};

export async function fetchPage({ url }) {
  const res = await fetch(
    `https://api.unblockingapi.com/unblock?url=${encodeURIComponent(url)}&render=true`,
    { headers: { "X-Api-Key": process.env.UNBLOCKINGAPI_KEY } }
  );
  const data = await res.json();
  return htmlToText(data.response);
}

Your agent loop passes fetchPageTool to the model, and when the model replies with a request to call fetch_page, you run fetchPage and send the result back as the tool's output.

Clean the page for the model

Rendered HTML is mostly scripts, styles and navigation. Sending all of it wastes tokens and buries the content. A small cleaner using Cheerio keeps the visible text and caps the length:

import * as cheerio from "cheerio";

export function htmlToText(html, maxChars = 12000) {
  const $ = cheerio.load(html);
  $("script, style, noscript, svg, nav, footer, iframe").remove();

  const text = ($("main").text() || $("body").text())
    .replace(/\s+/g, " ")
    .trim();

  return text.length > maxChars ? text.slice(0, maxChars) + " …[truncated]" : text;
}

Keep the source URL next to the text you return so the model - and your users - can cite where an answer came from.

Structured output

When the model only needs a few fields - a price, a title, a list of items - pass those instead of the whole page. You can build the extraction without writing selectors: open the page in the visual editor, tap the fields you want and save a template that returns clean JSON on every call. Ready-made options are in the templates library.

Keep RAG sources fresh

A retrieval index drifts out of date as the pages behind it change. A refresh job fetches each source on a schedule, hashes the cleaned text and only re-embeds pages whose hash changed:

import { createHash } from "node:crypto";

export async function refreshSource(source, store) {
  const text = await fetchPage({ url: source.url });
  const hash = createHash("sha256").update(text).digest("hex");

  if (hash === source.lastHash) return { url: source.url, changed: false };

  await store.reEmbed(source.url, text);          // your vector store
  await store.save(source.url, { hash, fetchedAt: new Date().toISOString() });
  return { url: source.url, changed: true };
}

Each fetch is a normal request, so you control the cost by choosing the schedule. For a longer look at monitoring pages over time, see the competitor research use case.

Handle failures honestly

Some fetches will fail or return a page with nothing useful on it. The important thing is what your tool tells the model. If it returns an empty string, the model may carry on as if it had read something. Return an explicit error instead:

const text = await fetchPage({ url });

if (!text || text.length < 200) {
  return "ERROR: the page returned no readable content. Do not guess - tell the user you could not read it.";
}
return text;
An agent that says "I couldn't read that page" is far more useful than one that fills the gap with a plausible guess.

Assistants over MCP

If you'd rather not write a tool at all, the open-source MCP server gives an assistant the same ability out of the box. Setup guides: Claude, Cursor, VS Code and Zed. The wider picture is on the AI agents and LLM apps page.

Responsible use

Use UnblockingAPI to read specific public pages for live, per-request use. It is not for bulk-collecting content to train models. Follow applicable laws, the terms of the sites you fetch and our Acceptable Use Policy.

Frequently asked questions

Give the model a tool that fetches a URL and returns the page content. When the model needs current information it calls the tool, your code fetches the page through UnblockingAPI, and the cleaned content goes back into the conversation. Assistants like Claude and Cursor can do this over MCP without any code.

Many sites build their content with JavaScript, so a plain HTTP request returns an empty shell, and some refuse automated requests. Fetching through a real browser returns the rendered page instead.

Usually not. Raw HTML is mostly markup and scripts, which costs tokens and distracts the model. Strip scripts and styles and pass the visible text, or extract only the fields you need as JSON.

It gives the model the actual page to answer from, which reduces answers invented from stale or missing information. It is not a guarantee, so keep the source URL with the content and let users check it.

As often as the underlying content changes. Fast-moving pages like prices or news may need daily or hourly refreshes; documentation may only need weekly. Compare a content hash and re-embed only the pages that actually changed.

No. UnblockingAPI is built for fetching specific public pages for live, per-request use. Use it within our Acceptable Use Policy and the terms of the sites you fetch.

Give your agent a page to read.

Test the fetching layer before building your parser. Start with 500 free credits and no credit card. See pricing for plan details.

One successful request is one credit. No multipliers for rendering, proxies or retries.

Last updated October 2026