← All projects

Data Pipelines · Open source

AI Web Scraping Pipeline

This pipeline is the engine behind AIUseCaseHub, released as a configurable open-source project. Three logical agents run inside one Azure Functions app: discovery finds candidate pages, extraction turns them into structured records, and review validates fields, checks for duplicates with embeddings, and scores groundedness against the source before anything is stored. Changing the schema, sources, and prompts adapts it to other research domains such as customer stories, construction projects, or public-sector modernization.

My role
Creator, architect, and builder
Started
Updated
Visibility
Open Source
Built with
Python · Azure Functions · Cosmos DB · Microsoft Foundry · Microsoft Agent Framework · Bicep and azd
agents: discovery, extraction, review
3
records powering AIUseCaseHub
~4,000
Azure deployment with azd and Bicep
1-click
AI web scraping pipeline schematic
How discovered elements flow through the discovery, extraction, and review agents before landing in Cosmos DB for search.

Problem

Scraping with LLMs is easy to start and hard to trust. Without deduplication, groundedness checks, and review gates, you end up with a large dataset nobody can rely on.

Solution

A production-oriented pipeline with configurable sources and schema, conservative cost controls, groundedness scoring with a retry, and an explicit review step before records are stored.

Architecture

Discovery, extraction, and review agents hand off work through Cosmos DB containers that act as queues. Approved records land in a Records container with vector search for duplicate detection. Everything deploys with azd and Bicep on a Flex Consumption plan.

Lessons learned

Quality gates are the product. The model call is the easy part; deduplication, groundedness, and token cost tracking decide whether the data is worth using.

More screenshots (1)
AI web scraping pipeline overview
The three-agent overview with out-of-the-box features and the Azure deployment footprint.

agents · web-scraping · groundedness · open-source