Skip to content
 
 

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Bright Data Web Research Agent

A minimal OpenAI Agents SDK example for a deep-research agent that uses Bright Data APIs for deterministic web search and page fetching.

The demo asks a company/product/market question, searches the web, reads source pages, and returns schema-validated JSON with citations.

Setup

python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .
Copy-Item .env.example .env

Fill in .env:

OPENAI_API_KEY=...
BRIGHT_DATA_API_TOKEN=...
BRIGHT_DATA_SERP_ZONE=...
BRIGHT_DATA_UNLOCKER_ZONE=mcp_unlocker

mcp_unlocker is Bright Data's default Web Unlocker zone for direct API scraping. Set BRIGHT_DATA_SERP_ZONE only if you also want SERP discovery.

Run

.\.venv\Scripts\python.exe -m bright_research_agent.agent `
  "What is the market positioning of Perplexity's enterprise search product?"

Progress and tool-call logs are written to stderr so stdout remains valid JSON:

.\.venv\Scripts\python.exe -m bright_research_agent.agent `
  "What is the market positioning of Perplexity's enterprise search product?" `
  --log-level INFO

Set --log-level WARNING or LOG_LEVEL=WARNING for quieter output.

If the OpenAI request times out, give the model call more room and reduce turns:

.\.venv\Scripts\python.exe -m bright_research_agent.agent `
  "What is the current landscape of GTM engineering?" `
  --openai-timeout 300 `
  --max-turns 6

For deeper public LinkedIn/company intelligence, use LinkedIn mode and increase the source budget:

.\.venv\Scripts\python.exe -m bright_research_agent.agent `
  "What are the latest public signals about LinkedIn's B2B advertising product strategy?" `
  --mode linkedin `
  --max-sources 6 `
  --max-turns 8

To scrape a public LinkedIn URL directly through Bright Data Web Unlocker, pass --url. This skips SERP discovery and only requires BRIGHT_DATA_API_TOKEN plus the Unlocker zone:

.\.venv\Scripts\python.exe -m bright_research_agent.agent `
  "Summarize this public LinkedIn company profile with citations." `
  --url "https://www.linkedin.com/company/linkedin/" `
  --mode linkedin `
  --max-sources 1

The final output is JSON matching the Pydantic schema in src/bright_research_agent/schemas.py.

What This Demonstrates

  • SERP discovery through Bright Data SERP API.
  • Page retrieval through Bright Data Unlocker API.
  • OpenAI Agents SDK tool orchestration.
  • Pydantic output validation for citation-backed research JSON.

Notes

  • Treat scraped content as untrusted input. The agent instructions explicitly tell the model not to follow instructions found inside retrieved pages.
  • Keep max_sources low during demos so the workflow stays fast and inexpensive.
  • This API-first version is the clearest starting point for retries, concurrency, metrics, and cost controls.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages