How to Feed Any Page to a Neural Network Without Ad Clutter and Extra Tokens
pullmd converts any webpage to clean Markdown for LLM agents, stripping ads and navigation clutter while preserving content structure and saving tokens.
Language
HomeLanguages
Sections
pullmd converts any webpage to clean Markdown for LLM agents, stripping ads and navigation clutter while preserving content structure and saving tokens.
A look at fast-plate-ocr, a lightweight library for license plate recognition that achieves up to 14,000 plates per second using compact CCT models and ONNX Runtime.
A Node.js library that extracts structured JSON from messy PDFs and scans using AI, supporting both cloud and local vision models.
Tesseract.js brings OCR directly to the browser. No backend, no cloud payments — just pure client-side text recognition with WebAssembly.
Ever needed to extract data from a PDF or scanned document? Text-extract-api offers an elegant solution combining modern OCR and language models.
Discover PDF Craft – a powerful Python tool that converts scanned PDFs into editable Markdown and EPUB formats with smart OCR.