-

Reconstructing the Table of Contents a PDF Forgot to Ship, So RAG Can Scope by Section
Large Language ModelsEnterprise Document Intelligence [Vol.1 #5septies] – When a PDF prints a contents page but exposes…
14 min read -

Enterprise Document Intelligence [Vol.1 #5sexies] – image_df tells you where every picture is. Turning the…
17 min read -

Parse Scanned PDFs for RAG with EasyOCR: Free OCR Gives You Words, Not a Document
Large Language ModelsEnterprise Document Intelligence [Vol.1 #5quinquies] – Same 1974 scanned PDF, two engines. EasyOCR recovers text.…
15 min read -

Enterprise Document Intelligence [Vol.1 #5ter] – Table cells, OCR, captions, headings: cloud-grade structure, running on…
19 min read -

Enterprise Document Intelligence [Vol.1 #5A] – Document signals (metadata, native TOC, source software) and page-level…
23 min read -

I created a graph storage from dozens of annual reports (with tables)
5 min read -

Anthropic Claude 3.5 now understands PDF input
8 min read -

A step-by-step guide to get underlined text as an array from PDF files.
5 min read -

Leveraging zero-shot labeling
5 min read
