What this tool does
- Finds columns automatically — column boundaries are inferred from where the text actually sits on the page
- One sheet per page — or everything merged into a single sheet, your choice
- XLSX or CSV — a real Excel workbook, or plain CSV for scripts and databases
- Numbers stay numbers — currency symbols, thousands separators and bracketed negatives are parsed into numeric cells you can sum
- Editable preview — see the detected grid and adjust the column sensitivity before exporting
How column detection works
A PDF table has no rows and columns in it — only text positioned on a page, which is why copying a table out of a PDF usually produces mush. This tool reads the x and y coordinate of every text fragment, groups fragments sharing a baseline into rows, then looks for the vertical gaps that persist across those rows. Gaps that line up down the page are column boundaries; gaps that do not are just spaces between words.
That works well on the kind of table PDFs are full of — bank statements, invoices, price lists, financial reports. It struggles with merged cells, cells containing wrapped multi-line text, and tables whose columns are separated by ruled lines rather than whitespace. The sensitivity slider lets you widen or narrow the gap threshold when a column is being split or merged wrongly.
Why extract tables
- Get a bank or card statement into a spreadsheet for reconciliation
- Turn a supplier's PDF price list into something you can sort and filter
- Pull figures out of a published report to chart them
- Feed invoice line items into accounting software
Frequently asked questions
Are my PDFs uploaded anywhere?
No — the file is opened and rewritten entirely inside your browser using JavaScript. Nothing is uploaded to a server, so your document never leaves your device.
Why are my columns split in the wrong place?
The gap threshold is too low or too high for your table. Move the column sensitivity slider and the preview updates immediately.
Does it work on scanned PDFs?
No — there are no text coordinates in an image. OCR the file first, though be aware OCR output is less precisely positioned, so table detection will be rougher.
Will formulas come across?
No. A PDF only ever contains the calculated values — the formulas were lost when the file was created, not when it was converted.
XLSX or CSV?
XLSX if you want multiple sheets and typed numeric cells. CSV if something downstream needs to parse it, keeping in mind CSV holds only one sheet.