Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
0.5.1Bright Data is a web data platform that provides proxies, scrapers, and structured data feeds at scale. This Arcade toolkit enables search, scraping, and structured data extraction from the web without bot detection blocking requests.
Capabilities
- Web scraping: Fetch any public webpage and return its content as clean Markdown.
- Multi-engine search: Query Google, Bing, or Yandex with configurable parameters including result count, search type (web/images), and country targeting.
- Structured data feeds: Extract pre-parsed, schema'd data from major platforms — Amazon (products, reviews), LinkedIn (people, companies), Instagram, Facebook, X, YouTube, Zillow, Booking.com, and ZoomInfo — without writing custom parsers.
Secrets
This toolkit requires two secrets to authenticate with Bright Data.
-
BRIGHTDATA_API_KEY— Your Bright Data account API key. Obtain it from the Bright Data dashboard under Account Settings → API Token. Any paid or trial Bright Data account can generate one. -
BRIGHTDATA_ZONE— A Bright Data zone identifier that determines which proxy network or dataset product is used for requests. Zones are created and managed in the Bright Data control panel under Proxies & Scraping Infrastructure. Each zone corresponds to a specific product (e.g., Web Unlocker, Scraping Browser, or a dataset feed). Use the zone name (not the full connection string) as the secret value.
For general guidance on configuring secrets in Arcade, see the Arcade secrets docs. Secrets can also be managed at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MADE UP LINKS - IF LINKS ARE NEEDED, EXECUTE search_engine FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 2 |