Skip to main content

Command Palette

Search for a command to run...

Web Page Content Readability API Documentation

Published
6 min readView as Markdown
Web Page Content Readability API Documentation

Webpage Readable Content Extraction API

Category: Web Tools
Endpoint: /v1/websitetools/readability
Method: POST
Description: Intelligently extracts key elements from web articles to deliver refined, reader-focused content.

Overview

The Webpage Readable Content Extraction API offered by GuGuData provides an efficient and precise way to parse, clean, and extract essential article information from any web page. This API streamlines the process of retrieving the title, author, excerpt, main text, language, and other critical metadata. Whether you're looking to enhance SEO, perform text analytics, or build content-driven applications, this API enables you to focus on the core readable sections of an article without distractions such as ads, scripts, or complex HTML structures.

With second-level parsing performance and nationwide multi-node CDN deployment, the API ensures both reliability and high concurrency support, making it suitable for enterprise-grade text analysis, automated news aggregation, or content curation tools.

Key advantages include:

  • Intelligent Extraction: Automatically identifies and retrieves main article content, including metadata such as publication time and byline.
  • HTML or URL Input: Submit raw HTML directly or specify a webpage URL for parsing.
  • Various Elements Retrieval: Get everything from article title and content to excerpt, language, and website name.
  • High Concurrency: Handle large volumes of requests with minimal latency.
  • HTTPS/TLS Ready: End-to-end encryption and compatibility with Apple ATS standards.

Request URL

POST https://api.gugudata.io/v1/websitetools/readability

Demo URL

Try out our live demo endpoint to see how the Readability API works:

https://api.gugudata.io/v1/websitetools/readability/demo

Request Parameters

Below are the key parameters needed to use the Readability API effectively:

ParameterTypeRequiredDefault ValueDescription
appkeystringYesYOUR_APPKEYThe APPKEY obtained from GuGuData's Developer Center.
htmlstringNoYOUR_VALUEOptional. Raw HTML content of the webpage you want to process. You must provide either html or url.
urlstringNoYOUR_VALUEOptional. The URL of the webpage you want to process. You must provide either html or url. For pages with anti-crawling measures, content retrieval may fail.

Note: If both html and url are provided, the API will prioritize one parameter according to internal rules (generally taking html as primary).

Response Parameters

When your request is successful, you'll receive a JSON response containing metadata and extracted content. Below is a summary of the main fields:

ParameterTypeRequiredDescription
DataStatus.RequestParameterstringYesThe parameters you passed to the API.
DataStatus.StatusCodeintYesThe API response status code.
DataStatus.StatusDescriptionstringYesA brief message describing the status code.
DataStatus.ResponseDateTimestringYesThe timestamp when the data is returned.
DataStatus.DataTotalCountintYesTotal data count under this condition, mainly for pagination.
Data.TitlestringYesExtracted article title.
Data.BylinestringYesIdentified author or byline.
Data.DirstringYesDetected direction of the text (e.g., ltr, rtl).
Data.LangstringYesLanguage code of the extracted content.
Data.ContentstringYesThe refined HTML content of the article.
Data.TextContentstringYesPlain text of the article content, divided into paragraphs, without HTML tags.
Data.LengthintYesTotal length of the extracted content.
Data.ExcerptstringYesA short excerpt or summary of the article.
Data.SiteNamestringYesName of the website where the content was extracted.
Data.PublishedTimestring[]YesPotential article publication times, if found.

Error Codes

Below are common error codes you may encounter while using this API:

Error CodeError ContentNotes
200Normal returnRequest was successful.
400Parameter errorOne or more parameters are invalid or missing.
429Request frequency limitedExceeded the max allowed requests (100 per second).
403Account in arrearsYour account is suspended due to payment issues.
402APPKEY errorCheck if the APPKEY matches the one from the Developer Center.
500API response errorA general error indicating an unexpected issue with the API.

Features

  1. Intelligently extracts readable content – Eliminates unnecessary HTML, scripts, and ads.
  2. Provides HTML code of readable segments – Ensures your application can display or process content seamlessly.
  3. HTML or URL input – Offers flexibility for direct HTML submission or standard URL-based parsing.
  4. Retrieves multiple metadata points – Title, author, language, excerpt, published time, and more.
  5. Second-level parsing performance – Delivers rapid extraction, supporting high concurrency.
  6. HTTPS/TLS support – Fully compatible with TLS v1.0 to v1.3 and Apple ATS security standards.
  7. Nationwide multi-node CDN – Boosts response times and reliability with distributed infrastructure.
  8. Load-balanced architecture – Handles large volumes of requests without downtime.

Additional Notes for Developers

  • Choose the right parameter: If the source URL is behind heavy anti-crawling or requires a login, consider sending the cleaned html directly.
  • Focus on SEO & analytics: Extracted data can be used for advanced content analysis, SEO audits, or natural language processing.
  • High concurrency: The multi-node CDN and load balancing let you scale quickly and handle spikes in traffic.

Getting Started

  1. Sign up at GuGuData.io – Register to get your appkey.
  2. Send a POST request – Provide necessary parameters (html or url) along with appkey.
  3. Parse the response – Access the refined content and metadata in JSON format.
  4. Integrate – Incorporate this extracted data into CMS tools, analytics platforms, or mobile apps.

About GuGuData

GuGuData is a leading data solutions provider with nearly a decade of experience in delivering comprehensive APIs. Our offerings range from SSL Certificate Information to advanced Image Recognition functionalities. Throughout the years, we've maintained a commitment to organizing, cleaning, and integrating massive volumes of data, ensuring that developers, analysts, and businesses can effortlessly harness powerful data insights.

Key Highlights:

  • Nine years in business – Tested, resilient services trusted by global enterprises.
  • 4.2k+ APIs – A constantly growing library of specialized data and functionalities.
  • 95% happy customers – Recognized for reliability, performance, and innovation.

Whether you're extracting readable text for content creation, analyzing user engagement, or performing text-based research, GuGuData's Webpage Readable Content Extraction API is your gateway to consistent, high-quality data. Visit our official website to explore more APIs and start optimizing your web content today.

More from this blog

网页转 JSON API 怎么接:从 Prompt 到可校验的结构化数据

把网页内容变成 JSON,真正困难的部分通常不是“得到一段看起来像 JSON 的文本”,而是让字段含义、数据类型、缺失值、来源和更新规则都能被程序稳定处理。 语义化获取站点 JSON 结构内容 API 可以根据网页 URL 和自然语言 Prompt 提取自定义结构。它适合商品列表、文章索引、表格内容、站点研究和自动化数据准备等场景,但接口返回成功并不代表所有字段都已经过事实验证。生产接入仍需要本地

Aug 27, 20265 min read

地址逆编码接口 API

此文章对开放数据接口 API 之「地址逆编码接口 API」进行了功能介绍、使用场景介绍以及调用方法的说明,供用户在使用数据接口时参考之用,并且在目前更新的微信小程序实战开发项目中的使用场景。 1. 产品功能 此次开放了精准的地址坐标逆编码在线接口,用于对提供的 GPS 坐标转换为文字地址信息。 提供精准、高效的地理坐标逆编码接口; 返回的地址包含详细的位置信息; 一次可返回坐标周边的 10

Aug 3, 20261 min read

中英文排版规范化 API

此文章对开放数据接口 API 之「中英文排版规范化 API」进行了功能介绍、使用场景介绍以及调用方法的说明,供用户在使用数据接口时参考之用。 1. 产品功能 此次开放了中英文排版规范化在线接口,用于自动中英文排版、标点符号格式化,中英混排格式化 / 标点修正。 支持中英文混排格式化; 自动在汉字与英文字符、英文标点、数字间添加空格; 中文标点符号自动规范化,遵从 [标点符号用法 GB/T

Aug 3, 20261 min read

为阿里云站点部署免费 HTTPS

本文记录了部署在阿里云的站点,在申请了免费的 SSL 证书后如何正确的部署到站点上,让站点支持 HTTPS 访问。 阿里云引入了沃通作为 CA 证书供应商,并开放了免费 SSL 申请的页面,之前一直想给 咕咕监控 部署上全站 HTTPS,所以就申请了一个,但是部署的过程中遇到了些问题,所以记录下来备忘。 1. 证书申请 在阿里云后台的 CA 管理页面,填写相关的信息后就可以申请到一张免费的 CA

Aug 3, 20261 min read

GuGuData

178 posts