Extract topics and full post text from any Discourse forum. Clean plain-text output for LLM training, RAG corpora and community research.
1
Users / 30d
5
Runs / 30d
—
Rating
Updated 15 Sept 2026
Apify builder
@bindler
Leaderboard position
out of 6,249
New Actors / 30d
#322
Total Actors
#1217
Active users / 30d
#1655
Total users
#2839
Runs / 30d
#3571
Total runs
#5311
Portfolio stats
New Actors / 30d
7
Total Actors
7
Active users / 30d
7
+5 since first snapshot
Total users
14
Runs / 30d
32
+27 since first snapshot
Total runs
44
Portfolio history
Daily publishing
7
Each bar represents one UTC day.
Actor portfolio
7 matching of 7 Actors · page 1 of 1
Extract topics and full post text from any Discourse forum. Clean plain-text output for LLM training, RAG corpora and community research.
1
Users / 30d
5
Runs / 30d
—
Rating
Scrape any documentation site to clean markdown. Works on Docusaurus, Mintlify, GitBook, MkDocs, ReadTheDocs and more. Preserves code blocks for RAG and LLM training.
1
Users / 30d
5
Runs / 30d
Turn any public GitHub repository into a clean dataset of code and documentation files. One download, no API token, no rate limits. Built for AI coding assistants and RAG.
1
Users / 30d
5
Runs / 30d
Convert PDFs to clean markdown with real reading order. Handles two-column layouts, detects headings and paragraphs. Built for RAG pipelines and LLM ingestion.
1
Users / 30d
4
Runs / 30d
Extract every table from any web page into clean rows, JSON and markdown. Correctly handles colspan, rowspan and stacked headers that break other extractors.
1
Users / 30d
3
Runs / 30d
Scrape posts, pages and content from any WordPress site via the REST API. Clean plain text with authors, categories and tags, ready for LLM training and RAG.
1
Users / 30d
5
Runs / 30d
Scrape all articles from any Zendesk Help Center. Full plain-text bodies with sections, categories and labels. Ready for RAG, AI support bots and knowledge base migration.
1
Users / 30d
5
Runs / 30d
—
Rating
—
Rating
—
Rating
—
Rating
—
Rating
—
Rating