Skip to content

About

YouTube Video Scraper retrieves a video's title, description, channel, engagement counts, chapters, related videos, and available comment or transcript follow-up links through SerpApi's `youtube_video` engine.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

YouTube Video Scraper

YouTube Video Scraper tool

YouTube Video Scraper retrieves a video's title, description, channel, engagement counts, chapters, related videos, and available comment or transcript follow-up links through SerpApi's youtube_video engine.

Get structured JSON for applications or Markdown for LLMs and AI agents without maintaining page parsers or proxies. This API retrieves video-page data, not downloadable video or audio files. Comments, replies, and transcript text require separate requests.

How to scrape a YouTube video?

Send a GET request using the video ID from youtube.com/watch?v=video_id or youtu.be/video_id:

https://serpapi.com/search?engine=youtube_video&v=vFcS080VYQ0&gl=us&hl=en&api_key=YOUR_SERPAPI_API_KEY

Code examples

cURL integration: JSON

JSON is the default, so no output parameter is needed:

curl --get https://serpapi.com/search \
 -d engine="youtube_video" \
 -d v="vFcS080VYQ0" \
 -d api_key="YOUR_SERPAPI_API_KEY"

These minimal commands print the API response directly, including any error message. For values containing spaces or special characters, use --data-urlencode instead of -d. The Python and JavaScript examples below include error handling for automation.

Output formats: JSON and Markdown

JSON is the default and preserves fields and continuation tokens for programmatic workflows. Set output=md for Markdown text optimized for LLMs and AI agents. Markdown conveys most of the same information, but its layout is not a guaranteed substitute for the JSON schema.

The official parameter table also accepts output=html. However, its HTML Results section explicitly says this engine has no full HTML response, only text. Do not expect a complete YouTube webpage; search_metadata.prettify_html_file, when returned, links to a prettified result.

Markdown request

curl --get https://serpapi.com/search \
 -d engine="youtube_video" \
 -d v="vFcS080VYQ0" \
 -d output="md" \
 -d api_key="YOUR_SERPAPI_API_KEY"

Read successful Markdown responses as text, not with a JSON parser or getJson. Inspect error responses rather than treating them as video content.

Python integration

Install requests and save the following code as main.py:

pip install requests
import requests

SERPAPI_API_KEY = "YOUR_SERPAPI_API_KEY"
params = {
    "api_key": SERPAPI_API_KEY,
    "engine": "youtube_video",
    "v": "vFcS080VYQ0",
    "gl": "us",
    "hl": "en",
    "output": "json",
}

search = requests.get("https://serpapi.com/search", params=params, timeout=60)
search.raise_for_status()
data = search.json()
if "error" in data:
    raise RuntimeError(data["error"])
print(data)

For Markdown, keep the imports, API key, and parameter definitions above, then replace the request and response-handling lines with:

params["output"] = "md"
search = requests.get("https://serpapi.com/search", params=params, timeout=60)
search.raise_for_status()
print(search.text)

JavaScript integration: JSON

Install the SerpApi JavaScript package and save this CommonJS example as index.cjs:

npm install serpapi
const { getJson } = require("serpapi");

async function main() {
  const data = await getJson({
    api_key: "YOUR_SERPAPI_API_KEY",
    engine: "youtube_video",
    v: "vFcS080VYQ0",
    gl: "us",
    hl: "en",
    timeout: 60000
  });
  if ("error" in data) {
    throw new Error(data.error);
  }
  console.log(data);
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

getJson selects JSON automatically. Its client-side timeout is in milliseconds, not an engine parameter. The promise rejects on HTTP, network, or timeout failures; the explicit check handles a JSON API error.

JavaScript integration: Markdown

This standalone alternative requires Node.js 18 or later and no packages. Save it as markdown.cjs:

async function main() {
  const params = new URLSearchParams({
    api_key: "YOUR_SERPAPI_API_KEY",
    engine: "youtube_video",
    v: "vFcS080VYQ0",
    gl: "us",
    hl: "en",
    output: "md"
  });
  const response = await fetch(`https://serpapi.com/search?${params}`, {
    signal: AbortSignal.timeout(60000)
  });
  const text = await response.text();
  if (!response.ok) {
    throw new Error(`SerpApi HTTP ${response.status}: ${text}`);
  }
  console.log(text);
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

Other programming languages

Use an HTTP GET client with encoded parameters, or explore SerpApi Integrations.

YouTube Video Scraper parameters

Name Description Requirement
engine Must be youtube_video. Required
api_key Your private SerpApi API key. Required
v YouTube video ID, not a search query or full URL. Supply it for the initial video lookup. The official parameter table labels it optional; this guide includes it in initial and continuation requests. Optional in docs; used here
gl Two-letter country code, such as us, uk, or fr. See Google countries. Optional
hl Two- or three-letter language code, optionally with a region, such as en, en-gb, or es-419. See Google languages. This localizes video-page data, not transcript-track selection. Optional
next_page_token A returned token for related videos, comments, comment sorting, or replies; see the workflow below. Optional
output json (default), md, or html. HTML output is text rather than a full HTML page for this engine. Optional
no_cache false by default; true bypasses the cache. Exact matching queries and parameters can use cached results for one hour; cached searches are free and do not count toward monthly searches. Do not combine with async. Optional
async false by default. true submits a search for later retrieval through the Search Archive API, using search_metadata.id. Do not combine with no_cache or use with Ludicrous Speed enabled. Optional
zero_trace Enterprise-only ZeroTrace mode; false by default. true skips storing search parameters, files, and metadata, making debugging harder. Optional
json_restrictor Select JSON fields to reduce payload size. Keep any tokens required by your workflow. See JSON Restrictor. Optional

See the official YouTube Video API documentation for the current contract. There is no documented q, numeric page, start offset, or configurable comment page size for this engine.

Related videos, comments, sorting, and replies

The initial video response contains metadata and, when available, continuation tokens. Comments and replies are not included in the initial response. Each continuation is a separate API request.

What to retrieve Token to put in next_page_token Follow-up result
More related videos Top-level related_videos_next_page_token related_videos; continue with the new related_videos_next_page_token.
Initial or subsequent comments Top-level comments_next_page_token comments; continue with the new comments_next_page_token.
A particular comment order comments_sorting_token[i].token, selected using that entry's title comments in the chosen order, then follow comments_next_page_token.
Replies to one comment comments[i].replies_next_page_token replies and comment_parent_id; continue using the top-level replies_next_page_token from that reply response.

Keep engine=youtube_video, the video's v, your localization, API key, and output=json. Send one token per request, replacing the previous next_page_token. Treat tokens as opaque and URL-encode them through your client; do not construct, trim, or reuse tokens from unrelated videos. Do not mix comment and reply pagination state.

Stop when the relevant token is absent. Page lengths vary; a short page is not a reliable end signal. Use a request budget and guard against repeated tokens.

For example, append this bounded continuation example to the Python JSON example, retaining its initial data and params. It retrieves at most three comment pages:

token = data.get("comments_next_page_token")
seen_tokens = set()

for _ in range(3):
    if not token:
        break
    if token in seen_tokens:
        raise RuntimeError("Repeated comment pagination token")
    seen_tokens.add(token)
    page_params = {**params, "output": "json", "next_page_token": token}
    search = requests.get(
        "https://serpapi.com/search", params=page_params, timeout=60
    )
    search.raise_for_status()
    page = search.json()
    if "error" in page:
        raise RuntimeError(page["error"])
    print(page.get("comments", []))
    token = page.get("comments_next_page_token")

For a specific ordering, initialize token from the desired entry in data["comments_sorting_token"] instead. Choose from returned titles rather than assuming an array position.

Transcript follow-up

transcript.serpapi_link, when available, leads to the separate youtube_video_transcript engine. The video response does not embed the transcript segments. Use the returned link's parameters and your private API key, or request the transcript engine directly with the video ID. See the YouTube Video Transcript Scraper for track selection and millisecond timestamps.

Available data on YouTube videos (JSON response)

The following is a field guide, not a literal API response. Fields depend on what YouTube exposes and which request you make:

{
  "title": "String: video title",
  "thumbnail": "String: video thumbnail URL",
  "channel": {
    "name": "String: channel name",
    "link": "String: channel URL",
    "thumbnail": "String: channel thumbnail URL",
    "subscribers": "String: display subscriber count",
    "extracted_subscribers": "Integer: numeric subscriber count when returned",
    "verified": "Boolean: channel verification"
  },
  "views": "String: display view count",
  "extracted_views": "Integer: numeric view count",
  "likes": "String: display like count",
  "extracted_likes": "Integer: numeric like count",
  "live": "Boolean: live video indicator",
  "published_date": "String: publication date",
  "description": {
    "content": "String: description in plain text",
    "links": [
      {
        "start_index": "Integer: link text start index in content",
        "length": "Integer: link text length",
        "text": "String: link label",
        "url": "String: destination URL",
        "type": "String: link type when available"
      }
    ]
  },
  "chapters": [
    {
      "title": "String: chapter title",
      "thumbnail": "String: chapter thumbnail URL",
      "time_start": "Integer: chapter start in seconds"
    }
  ],
  "transcript": {
    "serpapi_link": "String: transcript API follow-up URL"
  }
}
Additional field What it contains
search_metadata Request ID, status, timing, and response-file URLs when available. search_metadata.status indicates Processing, Success, or Error; top-level error contains a failure message.
search_parameters Parameters associated with the search.
related_videos[] video_id or playlist_id, link, serpapi_link, title, publication date, view counts, length, channel, badges in extensions, and thumbnail.static / thumbnail.rich. Some entries are playlists or live content.
end_screen_videos[] Suggested end-screen videos with IDs, links, titles, channel, views, length, and a thumbnail URL string.
additional_info[] Rows with title, text, and link.
shopping_results[] Product title, description, thumbnail, price, extracted_price, vendor, and link.
comments[] Follow-up only: comment_id, link, channel, published_date, content, vote_count, extracted_vote_count, reply_count, and possible replies_next_page_token.
comment_count, extracted_comment_count Display and numeric comment counts, when returned. These are not necessarily the number collected.
replies[], comment_parent_id Reply follow-up only: comments-shaped reply entries and their parent comment ID.
Continuation and sorting tokens The exact keys described in the pagination table above. Not guaranteed to exist on every response.

Use extracted numeric counts for arithmetic, but preserve display strings and optionality. A missing field is not necessarily zero. Video chapters use seconds in time_start; the transcript engine uses milliseconds. This API does not document a top-level video-duration field; length is documented on related/end-screen video entries.

Use cases

  • Enrich a discovered video catalog with descriptions, channels, chapters, and public engagement counts.
  • Analyze comment themes or sentiment, keeping replies linked to their parent and respecting pagination limits.
  • Compare related-video recommendations across localized requests.
  • Build an AI research assistant using Markdown for context and JSON for reliable IDs, counts, and follow-up tokens.
  • Connect video metadata with timestamped transcripts for searchable content.

Blog tutorial

The tutorial is supplementary; use the current API documentation above for response shapes and pagination. In particular, do not assume a fixed number of comments per page.

Video tutorial

Contacts

Feel free to reach out via contact@serpapi.com.

About

YouTube Video Scraper retrieves a video's title, description, channel, engagement counts, chapters, related videos, and available comment or transcript follow-up links through SerpApi's `youtube_video` engine.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors