qocgpiywhm
API Reference

Web Content Cleaner

POST
/api/v1/tools/clean_web_content

Use to extract the main article or document text from a single-page URL. Not for homepages or hubs (Web Index Cleaner) and not a link inventory (Link Extractor). Server HTML only; JavaScript is not run. Output is text or markdown. If the requested extract has no main content, the other profile is tried once in the same request. The HTTP body has correlation_id and data. The item array is data.results.

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

const body = JSON.stringify({  "url": "https://en.wikipedia.org/wiki/HTTP_402",  "output_format": "text",  "include_title": true})fetch("https://example.com/api/v1/tools/clean_web_content", {  method: "POST",  headers: {    "Content-Type": "application/json"  },  body})
{  "correlation_id": "00000000-0000-4000-8000-000000000001",  "data": {    "results": [      {        "data": {          "title": "List of HTTP status codes - Wikipedia",          "content": "This article lists standard and notable non-standard HTTP response status codes. 402 Payment Required is reserved for future use.",          "word_count": 20,          "format": "text"        },        "meta": {          "url": "https://en.wikipedia.org/wiki/HTTP_402",          "final_url": "https://en.wikipedia.org/wiki/HTTP_402",          "status_code": 200,          "content_type": "text/html; charset=UTF-8",          "truncated": false,          "javascript_rendered": false,          "via_proxy": false,          "retries": 0,          "extraction_profile_requested": "content",          "extraction_profile": "content",          "fallback_used": false        }      }    ]  }}