DocuFlow
Unstructured web pages, organized into useful data.
A distributed crawling and content-processing project that extracts page-level information, stores structured documents, and exposes search and retrieval APIs.
The problem
Website content is easy to browse but difficult to reuse systematically. Repeat crawls also need to distinguish changes from duplicate work.
How it comes together
Crawling, extraction, cleanup, and storage are separated into a pipeline. Each page becomes a structured MongoDB document with content, links, and metadata that the API can retrieve.
Follow the flow.
Page-level extraction of text, images, links, and metadata.
Incremental updates and deduplication.
Distributed processing and containerized services.
FastAPI interfaces for search, retrieval, and analytics.
Why this approach?
Store content at page granularity with a consistent schema, preserving the relationship between source URLs and extracted material.