Privacy-first semantic search for scholars who work with serious text collections. Your books, your metadata, your machine.
RAG — retrieval-augmented generation — is how an AI reasons over your own texts: it retrieves passages and feeds them to the model. But most RAG is naive: everything goes in as chunks, everything comes out as chunks. A footnote weighs the same as an abstract; tag schemas, cross-references, tables of contents, chapter boundaries — invisible. Generic models stay generic precisely where you need them sharpest: in your own field. While the RAG market leader just reversed course and now compiles structure upstream of retrieval, Archilles uses what’s already there.
A scholar’s library carries two kinds of structure already. Outside each document: the tags, highlights, cross-references and reading notes that took a decade to settle. Inside each document: the table of contents, chapters and headings — the architecture of the text itself. Archilles indexes both — structure-aware chunking on the inside, your curation on the outside — and hands the model the edges naive retrieval cannot recover. The hardest part of intelligent search is intelligence. Scholars bring their own.
No telemetry, no tracking, no upload you did not choose. The reasoning goes to the model you pick — a cloud assistant by deliberate decision, or a local one. Full sovereignty over research data that took you years to build.
Results cite page, chapter and section. In volumes Scriptor has prepared, the page is the one the book prints, and a page that was inferred rather than read off the page says so. Any citation can be checked later: does that page still carry that wording?
Runs natively in clients that launch local MCP servers: Claude Desktop, Claude Code, Cursor, Codex CLI, Windsurf, Cline, Cherry Studio. ChatGPT and Gemini connect through a bridge*.
Not another single-app tool: Archilles brings your reference manager, your vault and your e-book library into one index — searched with one query, cited to page and chapter. Adapters auto-detect your layout; over thirty file types; no migration.
Search German, English, Latin, Greek and French together. BGE-M3 embeddings understand meaning across languages.
Custom fields, reading status, project tags, annotations. The organising work you already did becomes a research superpower.
* ChatGPT and Gemini accept only a public HTTPS address, which in practice means a tunnel. Archilles speaks stdio, SSE and Streamable HTTP; the MCP guide documents the working setup. ↩
Set a library path to your Zotero, Obsidian vault, Calibre folder, or any directory.
Batch-index by tag, author, or the whole library. GPU-accelerated, resumable, crash-safe.
Hybrid semantic plus keyword search with RRF fusion and optional cross-encoder reranking.
Results with page numbers and chapters. Export to BibTeX, RIS, EndNote, JSON, CSV.
Find all discussions of trade routes between Mediterranean and Northern Europe before 1500.
Searches across Latin primary sources, German monographs, and English translations simultaneously.
Trace the motif of unreliable narrators across these fifty twentieth-century novels.
Finds passages that demonstrate the concept — even when the texts never name it.
Compare views on the hard problem of consciousness across Chalmers, Dennett and Nagel.
Precise name matching meets semantic understanding of the underlying concepts.
Find all precedents on liability for AI-generated decisions across my EU law collection.
Commentary, case law and regulatory texts searched in one pass. Custom fields act as filters.
Archilles is open source and in active development. The code lives on GitHub. Field notes and longer-form thinking on Substack and LinkedIn.