PDFs and common images
PDF; JPG/JPEG, PNG, GIF, TIF/TIFF, WebP, HEIC/HEIF, BMP, AVIF, and SVG. The file must still be valid for decoding and content extraction.
Text, markup, and code
TXT, HTML/HTM, Markdown, LOG, JSON/JSONL/NDJSON, YAML, XML, CSV, TSV, and supported source files including Swift, JavaScript, TypeScript, Python, Go, and Rust.
Modern Office and OpenDocument
Word: DOCX, DOCM, DOTX, DOTM. Excel: XLSX, XLSM, XLTX, XLTM. PowerPoint: PPTX, PPTM, POTX, POTM. OpenDocument: ODT, ODS, ODP.
Reading material and email
RTF rich text, EPUB ebooks, and EML email. Supported extraction does not imply complete parsing of every attachment, embedded object, or complex layout.
Scope and processing limits
This list does not include legacy binary DOC/XLS/PPT, audio/video files, or arbitrary folders. Password-protected, damaged, oversized, or extraction-limited files may not be analyzable. Contact support with the format and error message if needed.