Web Content Fetcher
Scrapes web pages and WeChat articles to produce clean, noise-free Markdown content for processing, translation, or archival.
The source repository doesn't declare a license. Check its terms before reusing the code.
Key features
- Automatic noise removal for headers, footers, and sidebars
- Crawl4ai integration for high-quality web scraping
- Image preservation with automated alt text mapping
- Specialized WeChat article fetching with metadata extraction
- UTF-8 encoded Markdown output ready for downstream processing
Use cases
- Converting online articles and documentation into clean Markdown for offline archival
- Automating the collection of web-based data for summarizing multiple sources at once
- Extracting content from WeChat Official Accounts for translation or analysis
FAQ
When should I use this skill in my Claude Code workflow?
Use it when you need to ingest online documentation, blog posts, or WeChat articles for summarization, translation, or archival without the clutter of a standard copy-paste.
Can I use this skill for batch processing multiple URLs?
Absolutely. The skill is designed to be scriptable, allowing you to loop through multiple URLs to fetch and convert entire libraries of web content into structured Markdown files automatically.
Does this skill support WeChat (微信公众号) articles?
Yes, it features a specialized fetcher designed to handle WeChat's unique structure, including lazy-loaded images and metadata extraction that standard scrapers often miss.
What does the Web Content Fetcher skill do?
This skill allows Claude to scrape web pages and WeChat articles, stripping away 'noise' like headers, footers, and ads to produce clean, UTF-8 encoded Markdown content optimized for AI processing.
How does it improve my AI coding and research process?
It automates data collection using high-quality tools like Crawl4ai. By providing clean Markdown, it reduces token usage and improves the accuracy of Claude’s downstream analysis or translation tasks.
Related skills
Crawl4AI
Scrapes websites, extracts structured data, and automates web data collection pipelines using the Crawl4AI library.
Web Scraping Data Collection412 ptsFirecrawl API
Extracts clean web content, crawls entire domains, and searches the web directly through the Firecrawl API using terminal commands.
Web Scraping Data Collection612 ptsFirecrawl Web Scraping
Scrapes web content, maps site structures, and extracts structured data using advanced crawling and search capabilities.
Web Scraping Data Collection612 pts