Data Pipelines · Open source
AI Web Scraping Pipeline
This pipeline is the engine behind AIUseCaseHub, released as a configurable open-source project. Three logical agents run inside one Azure Functions app: discovery finds candidate pages, extraction turns them into structured records, and review validates fields, checks for duplicates with embeddings, and scores groundedness against the source before anything is stored. Changing the schema, sources, and prompts adapts it to other research domains such as customer stories, construction projects, or public-sector modernization.
- My role
- Creator, architect, and builder
- Started
- Updated
- Visibility
- Open Source
- Built with
- Python · Azure Functions · Cosmos DB · Microsoft Foundry · Microsoft Agent Framework · Bicep and azd
- agents: discovery, extraction, review
- 3
- records powering AIUseCaseHub
- ~4,000
- Azure deployment with azd and Bicep
- 1-click

Problem
Scraping with LLMs is easy to start and hard to trust. Without deduplication, groundedness checks, and review gates, you end up with a large dataset nobody can rely on.
Solution
A production-oriented pipeline with configurable sources and schema, conservative cost controls, groundedness scoring with a retry, and an explicit review step before records are stored.
Architecture
Discovery, extraction, and review agents hand off work through Cosmos DB containers that act as queues. Approved records land in a Records container with vector search for duplicate detection. Everything deploys with azd and Bicep on a Flex Consumption plan.
Lessons learned
Quality gates are the product. The model call is the easy part; deduplication, groundedness, and token cost tracking decide whether the data is worth using.
More screenshots (1)

agents · web-scraping · groundedness · open-source