MCP tool reference
92 tools, one MCP connection. 11 of them — the Scraper Studio lifecycle and the account tools — are this fork’s own; the rest are Bright Data’s upstream server, unchanged.
Build, run and self-heal custom scrapers for sites no pre-built extractor covers. This fork’s core contribution.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
scraper_createthis forkBuild a brand-new custom scraper for any public page from a natural-language description, using Bright Data Scraper Studio. Returns a collector_id immediately — AI generation takes 5–10 minutes.
scraper_statusthis forkPoll build or heal progress for a collector once, returning the current status and step.
scraper_runthis forkRun a collector against a URL and return the structured records it collected.
scraper_healthis forkRewrite a collector’s extraction logic from a plain-language description of what went wrong — which field went empty, what the page looks like now.
scraper_approvethis forkApprove or reject a heal or build that is waiting on manual confirmation.
scraper_ensurethis forkThe full lifecycle in one call: reuse a known scraper or build one, run it, check the data (not just the status code), and if it’s wrong, heal it and verify the fix actually held — escalating to a one-time rebuild if it still doesn’t.
scraper_registry_listthis forkList every scraper this account has built through Studio: domain, collector ID, health, and how many times it’s been repaired.
Zones, balance and spend for the Bright Data account, plus two extra scrape output formats.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
zones_listthis forkList the proxy zones on this Bright Data account, with the type of each one. Zones are the account-level resources that scraping and browser requests are billed against.
budget_statusthis forkShow the Bright Data account balance and what each zone has cost so far. Use this to check there is credit left before starting work that consumes it.
scrape_screenshotthis forkCapture a page as a PNG image, going through Bright Data’s unlocker so it works on sites with bot detection.
scrape_metadatathis forkFetch a page and return its HTTP status code, response headers and body as structured JSON, rather than converted text — for inspecting redirects, content type, or reachability.
Retail and marketplace datasets for product intel.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_amazon_productupstreamQuickly read structured Amazon product data. Requires a valid product URL containing /dp/. Often faster and more reliable than scraping.
web_data_amazon_product_reviewsupstreamQuickly read structured Amazon product review data. Requires a valid product URL containing /dp/.
web_data_amazon_product_searchupstreamRetrieve structured Amazon search results. Requires a search keyword and Amazon domain URL; limited to the first page of results.
web_data_walmart_productupstreamQuickly read structured Walmart product data. Requires a product URL containing /ip/.
web_data_walmart_sellerupstreamQuickly read structured Walmart seller data. Requires a valid Walmart seller URL.
web_data_ebay_productupstreamQuickly read structured eBay product data. Requires a valid eBay product URL.
web_data_homedepot_productsupstreamQuickly read structured Home Depot product data. Requires a valid homedepot.com product URL.
web_data_zara_productsupstreamQuickly read structured Zara product data. Requires a valid Zara product URL.
web_data_etsy_productsupstreamQuickly read structured Etsy product data. Requires a valid Etsy product URL.
web_data_bestbuy_productsupstreamQuickly read structured Best Buy product data. Requires a valid Best Buy product URL.
web_data_google_shoppingupstreamQuickly read structured Google Shopping product data. Requires a valid Google Shopping product URL.
Social networks, UGC platforms, and creator insights.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_linkedin_person_profileupstreamQuickly read structured LinkedIn people profile data. Requires a valid LinkedIn profile URL.
web_data_linkedin_company_profileupstreamQuickly read structured LinkedIn company profile data. Requires a valid LinkedIn company URL.
web_data_linkedin_job_listingsupstreamQuickly read structured LinkedIn job listings data. Requires a valid LinkedIn jobs or search URL.
web_data_linkedin_postsupstreamQuickly read structured LinkedIn posts data. Requires a valid LinkedIn post URL.
web_data_linkedin_people_searchupstreamQuickly read structured LinkedIn people search data. Requires a LinkedIn people search URL.
list_dataset_fieldsupstreamList the filterable fields of a searchable dataset (field name, type, and description). Call this before search_dataset.
search_datasetupstreamSearch a Bright Data dataset by a filter and get matching records back directly — fast Elasticsearch-backed search, no trigger/poll cycle.
web_data_instagram_profilesupstreamQuickly read structured Instagram profile data. Requires a valid Instagram profile URL.
web_data_instagram_postsupstreamQuickly read structured Instagram post data. Requires a valid Instagram post URL.
web_data_instagram_reelsupstreamQuickly read structured Instagram reel data. Requires a valid Instagram reel URL.
web_data_instagram_commentsupstreamQuickly read structured Instagram comments data. Requires a valid Instagram URL.
web_data_facebook_postsupstreamQuickly read structured Facebook post data. Requires a valid Facebook post URL.
web_data_facebook_marketplace_listingsupstreamQuickly read structured Facebook Marketplace listing data. Requires a valid listing URL.
web_data_facebook_company_reviewsupstreamQuickly read structured Facebook company reviews data. Requires a company URL and review count.
web_data_facebook_eventsupstreamQuickly read structured Facebook events data. Requires a valid Facebook event URL.
web_data_tiktok_profilesupstreamQuickly read structured TikTok profile data. Requires a valid TikTok profile URL.
web_data_tiktok_postsupstreamQuickly read structured TikTok post data. Requires a valid TikTok post URL.
web_data_tiktok_shopupstreamQuickly read structured TikTok Shop product data. Requires a valid TikTok Shop product URL.
web_data_tiktok_commentsupstreamQuickly read structured TikTok comments data. Requires a valid TikTok video URL.
web_data_x_postsupstreamQuickly read structured X (Twitter) post data. Requires a valid X post URL.
web_data_x_profile_postsupstreamQuickly read structured X (Twitter) profile posts. Requires a valid X profile URL.
web_data_youtube_profilesupstreamQuickly read structured YouTube channel profile data. Requires a valid YouTube channel URL.
web_data_youtube_commentsupstreamQuickly read structured YouTube comments data. Requires a video URL and optional comment count.
web_data_youtube_videosupstreamQuickly read structured YouTube video metadata. Requires a valid YouTube video URL.
web_data_reddit_postsupstreamQuickly read structured Reddit post data. Requires a valid Reddit post URL.
web_data_reddit_commentsupstreamQuickly read structured Reddit comments data. Accepts an optional days_back parameter.
Bright Data Scraping Browser tools for live automation.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
scraping_browser_navigateupstreamOpen or reuse a scraping-browser session and navigate to the provided URL, resetting tracked network requests.
scraping_browser_go_backupstreamNavigate the active session back to the previous page and report the new URL and title.
scraping_browser_go_forwardupstreamNavigate the active session forward to the next page and report the new URL and title.
scraping_browser_snapshotupstreamCapture an ARIA snapshot of the current page listing interactive elements and their refs for later ref-based actions.
scraping_browser_fill_formupstreamFill multiple form fields in one call using refs from the latest ARIA snapshot.
scraping_browser_click_refupstreamClick an element using its ref from the latest ARIA snapshot; requires a ref and human-readable element description.
scraping_browser_type_refupstreamFill an element identified by ref from the ARIA snapshot, optionally pressing Enter to submit after typing.
scraping_browser_screenshotupstreamCapture a screenshot of the current page; supports optional full_page mode.
scraping_browser_network_requestsupstreamList the network requests recorded since page load with method, URL, and response status.
scraping_browser_wait_for_refupstreamWait until an element identified by ARIA ref becomes visible, with an optional timeout.
scraping_browser_get_textupstreamReturn the text content of the current page’s body element.
scraping_browser_get_htmlupstreamReturn the HTML content of the current page.
scraping_browser_scrollupstreamScroll to the bottom of the current page.
scraping_browser_scroll_to_refupstreamScroll the page until the element referenced in the ARIA snapshot is in view.
scraping_browser_select_refupstreamChoose an option in a dropdown by its visible label.
scraping_browser_check_refupstreamTick a checkbox or select a radio button. Does nothing if already ticked.
scraping_browser_uncheck_refupstreamUntick a checkbox. Does nothing if already unticked.
scraping_browser_hover_refupstreamMove the mouse over an element, to open hover menus or reveal content that only appears on hover.
scraping_browser_reloadupstreamReload the current page — useful after an action that changed server-side state, or to retry.
scraping_browser_cookiesupstreamList the cookies the current browser session holds, to check session or consent state.
scraping_browser_close_sessionupstreamClose the browser and end the session. The next browser tool call starts a fresh one.
Company, financial, and location intelligence datasets.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_yahoo_finance_businessupstreamQuickly read structured Yahoo Finance company profile data. Requires a valid Yahoo Finance business URL.
Company and location intelligence datasets.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_crunchbase_companyupstreamQuickly read structured Crunchbase company data. Requires a valid Crunchbase company URL.
web_data_zoominfo_company_profileupstreamQuickly read structured ZoomInfo company profile data. Requires a valid ZoomInfo company URL.
web_data_google_maps_reviewsupstreamQuickly read structured Google Maps reviews data. Requires a Maps URL and optional days_limit.
web_data_zillow_properties_listingupstreamQuickly read structured Zillow property listing data. Requires a valid Zillow listing URL.
web_data_booking_hotel_listingsupstreamQuickly read structured Booking.com hotel listing data. Requires a valid Booking.com listing URL.
App stores, news, and developer data feeds.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_github_repository_fileupstreamQuickly read structured GitHub repository file data. Requires a valid GitHub file URL.
web_data_reuter_newsupstreamQuickly read structured Reuters news data. Requires a valid Reuters article URL.
App store listing data.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_google_play_storeupstreamQuickly read structured Google Play Store app data. Requires a valid Play Store app URL.
web_data_apple_app_storeupstreamQuickly read structured Apple App Store app data. Requires a valid App Store app URL.
Travel booking information.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_booking_hotel_listingsupstreamQuickly read structured Booking.com hotel listing data. Requires a valid Booking.com listing URL.
Higher-throughput scraping utilities and batch helpers.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
search_engine_batchupstreamRun up to 10 search queries in parallel. Returns JSON for Google results and Markdown for Bing/Yandex.
scrape_batchupstreamScrape up to 10 webpages in one request and return an array of URL/content pairs in Markdown format.
scrape_as_htmlupstreamScrape a single webpage with advanced extraction and return the HTML response body.
scrape_screenshotthis forkCapture a page as a PNG image, going through Bright Data’s unlocker so it works on sites with bot detection.
scrape_metadatathis forkFetch a page and return its HTTP status code, response headers and body as structured JSON, rather than converted text — for inspecting redirects, content type, or reachability.
extractupstreamScrape a webpage as Markdown and convert it to structured JSON using AI sampling, with an optional custom extraction prompt.
session_statsupstreamReport how many times each tool has been called during the current MCP session.
Measure and analyze AI/LLM brand visibility and generative engine optimization.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_chatgpt_ai_insightsupstreamSend a prompt to ChatGPT and get back AI-generated insights: answer text, citations, recommendations, and markdown. Useful for GEO and LLM-as-a-judge.
web_data_grok_ai_insightsupstreamSend a prompt to Grok and get back AI-generated insights as structured markdown.
web_data_perplexity_ai_insightsupstreamSend a prompt to Perplexity and get back AI-generated insights as structured markdown.
Developer tools and package information datasets.
search_engineupstreamScrape search results from Google, Bing, or Yandex. Returns SERP results in JSON for Google and Markdown for Bing/Yandex; supports pagination with the cursor parameter.
scrape_as_markdownupstreamScrape a single webpage with advanced extraction and return Markdown. Uses Bright Data’s unlocker to handle bot protection and CAPTCHA.
discoverupstreamSearch the web and rank results by AI-driven relevance. Returns scored results with title, description, URL, and relevance score. Supports intent-based ranking, geo-targeting, date filtering, and keyword filtering.
web_data_npm_packageupstreamQuickly read structured npm package data including latest version, README, dependencies, and metadata.
web_data_pypi_packageupstreamQuickly read structured PyPI package data including latest version, README, dependencies, and metadata.