Set it up from scratch in five minutes, then read, create and edit PDF, DOCX, Excel and PowerPoint from any AI client.
Node 20 or newer, git, and about 750 MB of disk. No account, no API key, and nothing listening on the network.
git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git cd leanzero-mcp-doc-processor npm install pwd # copy this — the client config needs the absolute path
There is no build step — the server runs from source. Puppeteer downloads Chromium during install, which is the slow part and what renders create-pdf output.
claude mcp add doc-processor --scope user \ -- node /ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js
{
"mcpServers": {
"doc-processor": {
"command": "node",
"args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
"env": {
"DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents"
}
}
}
}printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| node src/index.js 2>/dev/null | tail -1
# A healthy server answers with all 17 tools.Full detail — prerequisites, where files land, OCR, every client, and what each failure means — is in part 1 of the manual.
Three deployments, one codebase. Most people want the first, and the manual documents all three end to end.
Recommended for almost everyone
Your MCP client launches the server as a child process. No port, no listener, no auth to manage — and documents land on your own disk, in your own project.
When something remote has to reach it
The HTTP transport under launchd or systemd, published through a Tailscale Funnel with per-tenant bearer tokens. This is the exact recipe our own Mac Studio runs.
To evaluate it in a minute
One instance of this exact repo on our Mac Studio. Grab a free key and point any MCP client at it. An evaluation surface — not somewhere to put anything confidential.
One `claude mcp add` command — stdio for a local build, or `--transport http` for the hosted endpoint.
One entry in `claude_desktop_config.json`. Use the absolute `node` path — a GUI app does not inherit your shell PATH.
The same `mcp.json` shape. This is also the bridge Forge apps reach through, so it is worth getting exactly right.
Any MCP client takes `command` + `args` + `env`. That is the whole contract.
CogniRunner brokers every call from inside Forge, so your AI provider never sees the URL or the token.
Every question this server raises, answered in order: install it, connect your client, run it as a service the way we do, use all 17 tools, read the code, wire it into Jira through CogniRunner, and fix it when it misbehaves. 37 sections.
Written against the server's own source and against the live deployment behind the hosted demo — the launchd, Tailscale and hardening steps are the configuration actually running, with the secrets removed.
Node 20 or newer, git, about 750 MB of disk, and five minutes. No API key, no account, no LeanZero anything — the server is MIT-licensed and runs entirely on your machine.
MCP Doc Processor is a Node.js program that speaks the Model Context Protocol over stdin/stdout. Your AI client (Claude Code, Claude Desktop, LM Studio, Cursor, Cline, Qwen Code — anything that reads an mcp.json) launches it as a child process and talks JSON-RPC to it. That is the whole deployment model for the local case: there is no daemon to install, no port to open, and nothing listening on the network.
| What | Version | Why | Check it |
|---|---|---|---|
| Node.js | 20 or newer (we run 24 in production) | package.json declares engines: { node: ">=20" }. The code uses ESM, top-level await and the built-in fetch. | node --version |
| npm | 10 or newer (ships with Node 20+) | Installs the dependency tree. | npm --version |
| git | any | Cloning the repo. A downloaded ZIP works too. | git --version |
| Disk | ~750 MB | Measured on a clean clone: 185 MB of node_modules, plus ~565 MB of Chrome and chrome-headless-shell that Puppeteer downloads into ~/.cache/puppeteer for create-pdf. The browser cache is shared between clones, so a second checkout costs only the 185 MB. | df -h . |
| An MCP client | any | Claude Code, Claude Desktop, LM Studio, Cursor, Cline, Qwen Code, Roo Code, or your own MCP host. | — |
read-doc only when OCR-ing a scanned (image-only) PDF — see "How do I turn on OCR" below — and fact-check, which always needs a bearer for the web-search MCP plus a Serper key. Text PDFs, DOCX, XLSX and PPTX are parsed locally with no network access at all.Windows, macOS and Linux all work. The examples below use macOS/Linux paths; on Windows use the same commands in PowerShell and write paths as C:\\Users\\you\\... with doubled backslashes inside JSON.
Clone, install, note the path — then prove it works before you touch a client config.
git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git then cd leanzero-mcp-doc-processor. The repository is public, so no GitHub account or token is needed.npm install. This pulls @modelcontextprotocol/sdk, docx, xlsx-js-style, pdf-parse, mammoth, pptxgenjs, marked, jszip and puppeteer. Puppeteer downloads Chrome into ~/.cache/puppeteer on install — that is the slow part and roughly 565 MB of the total, and it is what renders create-pdf output.pwd (macOS/Linux) or cd (Windows) and copy the result. Every client config below needs the absolute path to src/index.js; a relative path will not work because the client launches the process from its own working directory.git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git
cd leanzero-mcp-doc-processor
npm install
pwd # copy this — you need the absolute path belowPUPPETEER_SKIP_DOWNLOAD=1 npm install. Everything works except create-pdf, which will error when called. Add Chromium later with npx puppeteer browsers install chrome.There is no build step. The server is plain ESM JavaScript and runs from source — src/index.js is the entry point for the stdio transport, src/server.js for the HTTP one. Nothing is transpiled or bundled, which means you can read and edit any file in the repo and restart your client to pick it up.
Send a real MCP handshake down the pipe and count the tools. If you get 17 back, the server is genuinely working — before any client is involved.
MCP over stdio is line-delimited JSON-RPC. That means you can drive the server from a shell with echo and a pipe, with no client in the loop. This is the single most useful debugging trick with this server, and it is worth doing once now so you know what a healthy response looks like.
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| node src/index.js 2>/dev/null \
| tail -1 | python3 -c 'import json,sys; print(len(json.load(sys.stdin)["result"]["tools"]), "tools")'17 tools. If you get that, the server is installed correctly and every problem from here on is a client-configuration problem, not an install problem.npm install did not finish. cd into the repo root and re-run npm install.node --version and upgrade.printf above closes the pipe; if you typed the JSON by hand, press Ctrl-D.2>/dev/null above is what keeps them apart.npm run lint:no-console-log exits non-zero and names the file — but nothing runs it for you: there is no CI configured in the repository, so run it yourself before committing. Use the logger in src/utils/logger.js, which writes to stderr.By default, in the working directory the client launched the server from — which for a coding agent is your open project. DOC_OUTPUT_DIR pins it somewhere fixed instead.
This is the question people get wrong most often, so it is worth being precise. src/tools/utils.js resolves the output root in three steps, first match wins:
X-Output-Dir header) — only exists on the HTTP transport, and it is sandboxed under a server-side base directory. A remote tenant can organise files into subfolders but cannot write anywhere it likes on the host.env block of your client config.logs/ folder appears in whichever project your agent opened — observed on a clean install. Add it to that project's .gitignore, or set DOC_OUTPUT_DIR and launch from elsewhere.X-Output-Dir on the hosted demo organises files on our Mac Studio, and you retrieve them over the signed /files/download link the tool returns. If you want documents to appear in your own Finder, run the server locally over stdio. There is no configuration that makes a remote process write to your laptop.Inside the output root, the server enforces a docs/ folder and sorts documents into one of six category subfolders (contracts, technical, business, legal, meeting, research) chosen by the auto-categoriser in src/utils/categorizer.js. Pass enforceDocsFolder: false on a create call to opt out and write exactly where you asked.
{
"mcpServers": {
"doc-processor": {
"command": "node",
"args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
"env": {
"DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents"
}
}
}
}One environment variable, and you choose the provider: a cloud vision model, or a local one in LM Studio so no page image ever leaves your machine.
A PDF that was produced by a word processor contains a text layer, and pdf-parse reads it with no AI and no network. A PDF that was produced by a scanner is a bag of images: there is nothing to extract, and the only way to read it is to look at it. That is what the vision path in src/services/vision-service.js is for, and it is the only part of this server that ever calls out to a model.
| Variable | Default | What it does |
|---|---|---|
Z_AI_API_KEY | (none) | The key. ZAI_API_KEY and ANTHROPIC_AUTH_TOKEN are also accepted as fallbacks. Without one, OCR is simply off and everything else still works. |
Z_AI_MODE | (unset) | ZAI (also Z_AI or Z; PLATFORM_MODE is read as a fallback) selects the Z.AI coding-plan endpoint https://api.z.ai/api/coding/paas/v4/ — not the standard one. Set Z_AI_CODING_PLAN=false alongside it for https://api.z.ai/api/paas/v4/. Anything else falls back to BigModel (open.bigmodel.cn). |
Z_AI_BASE_URL | auto from Z_AI_MODE | Overrides the base URL outright. This is the local-OCR switch — point it at any OpenAI-compatible /v1/ endpoint. |
Z_AI_CODING_PLAN | (unset) | Only read when Z_AI_MODE selects Z.AI. Anything other than the literal false keeps the coding-plan endpoint; set it to false if your key is for the standard API. |
Z_AI_VISION_MODEL | glm-4.6v | Model name sent in the request body. |
Z_AI_TIMEOUT | 300000 | Request timeout in milliseconds. Vision on a long scan is slow; 5 minutes is deliberate. |
SKIP_TABLE_EXTRACTION | true | Set false to also run table reconstruction on page images. Slower, and worth it on financial scans. |
The vision client posts an OpenAI-shaped chat/completions request to Z_AI_BASE_URL. Any server that speaks that shape works — including LM Studio on your own machine. Load a vision-capable model in LM Studio, start its local server, and point the MCP at it. Page images then never leave the machine.
{
"mcpServers": {
"doc-processor": {
"command": "node",
"args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
"env": {
"Z_AI_BASE_URL": "http://localhost:1234/v1/",
"Z_AI_VISION_MODEL": "<the model id LM Studio shows>",
"Z_AI_API_KEY": "lm-studio"
}
}
}
}baseUrl + "chat/completions" with no separator, so http://localhost:1234/v1 (no slash) produces .../v1chat/completions and 404s.isConfigured() refuses any key containing the substring "api" or "key", so a literal your-api-key reads as not-configured rather than failing later with a confusing 401. If OCR says "not configured" while you are sure the variable is set, check the value for those substrings — lm-studio and a real token both pass.To be precise about what leaves the machine: OCR never touches the web-search MCP, but it is not automatically local either — its one outbound call goes to Z_AI_BASE_URL, which is a cloud endpoint unless you point it at LM Studio as shown above. The other outbound path is fact-check, which calls the sibling web-search MCP — see the tools part.
Four jobs: read something, make something, change something, or keep track of what you made. The decision table is in the manual.
Read a PDF, DOCX, Excel or PowerPoint — summary, in-depth, or one focused question. Local path, or a remote HTTPS URL.
Decide the right format and tone before creating anything. Returns a ready-to-use creation plan.
A styled, editable Word DOCX — headings, tables, headers, footers, page numbers, 8 presets.
A final, print-and-sign PDF from markdown. Rendered through headless Chromium.
A README, spec or runbook — with an optional generated table of contents and YAML frontmatter.
An XLSX workbook: multiple sheets, styled headers, and live formulas.
An editable deck — one slide per '## ' heading, with bullets, tables, charts and speaker notes.
Append or replace in a DOCX. Append patches the XML in place, so formatting and images survive.
Append rows, add a sheet, or replace one sheet's data — keeping the workbook's styles.
Preview, append slides, or replace a slide in an existing deck.
Verify a document's claims against the live web. Calls the web-search MCP per claim and cites sources.
Search the registry by category, tag and title — to find a document or catch a duplicate.
List the blueprints a create call can be validated against.
Document DNA — project-wide header, footer and style defaults, memories, and usage-driven evolution.
Learn, list and delete structural blueprints extracted from documents you liked.
Fingerprint a document and get told when its structure changes underneath you.
Trace provenance — which sources informed a document, and what was derived from it.
About 18,000 lines of plain ESM JavaScript, no build step, no framework. The file-by-file tour is in the manual.
`src/index.js` speaks stdio, `src/server.js` speaks HTTP, and both dispatch through `tool-registry.js`. The transports cannot drift on what the tools are or what they accept.
`pdf-parse`, `mammoth` and `xlsx` read; `docx`, `pptxgenjs`, `xlsx-js-style` and headless Chromium write. Nothing in the read path touches the network unless you enable OCR.
System defaults, then project `.document-dna.json`, then user `.document-user.json`, with anything explicit in the call winning. That is why documents come out looking like yours without anyone remembering to ask.
`getOutputRoot()` resolves the per-request header first (sandboxed), then `DOC_OUTPUT_DIR`, then the working directory. A hosted tenant can organise files but cannot escape its base.
One JSON envelope in each direction lets any authenticated HTTPS endpoint feed files in and take files out. CogniRunner is the reference implementation, not a special case.
Every created document is indexed with category, tags and lineage — which is what makes duplicate detection, provenance and drift monitoring possible at all.
A workflow rule that can open the PDF somebody attached, decide whether it says what it needs to say, and attach a generated report back to the issue.
The hosted bridge works with every AI provider — CogniRunner dials the MCP's HTTPS URL from Forge. The local path works only with LM Studio, but keeps the documents on your own machine.
Compare them →CogniRunner mints a one-shot URL and bearer per attachment. The model calls read-doc with url + authHeader; the capability dies on first use, and nothing is copied to a shared disk.
Every MCP call is brokered by CogniRunner inside Atlassian Forge. The AI provider receives the tool's result — never the URL, never the bearer, never the raw attachment.
See CogniRunner →*.ts.net Tailscale Funnel URL on port 443 and nothing else — 8443 and 10000 are blocked, and an arbitrary domain is refused. Publish on a path prefix instead: tailscale funnel --bg --set-path=/docproc 10000. Full walkthrough →One instance of this exact repo, on our Mac Studio behind a Tailscale Funnel. Endpoint: https://worksmacstudio.tailfc4700.ts.net/docproc/mcp
This server is one half of a pair, and both are wired into the same Jira app.
The sibling server. fact-check calls it per claim, so if you want document verification you need both — same setup shape, same launchd and Tailscale recipe.
The Jira workflow app that consumes this MCP — AI validators, conditions and post-functions, with the MCP setup documented on its own page.
Everything we have written about running models on your own hardware — including the LM Studio setup this server's local OCR path uses.
If you would rather have this wired into your own stack than do it yourself — that is the work we do.
Mostly not. Fifteen of the seventeen tools need no key at all — PDF, DOCX, XLSX and PPTX are parsed locally with no network access. Two do: read-doc needs a vision key only to OCR a scanned, image-only PDF, which you can point at a local model in LM Studio instead of a cloud provider; and fact-check always needs a bearer for the web-search MCP plus your own Serper key.
In the working directory the client launched the server from, which for a coding agent is your open project. Set DOC_OUTPUT_DIR to pin a fixed location. A remote server writes to the remote host's disk — that is a client/server boundary, not a setting, so run it locally over stdio if you want documents in your own Finder.
Yes, MIT licensed, full source on GitHub. Clone it, run it, fork it, ship it in a product. The hosted demo is one deployment of that same repo, not a different build.
The code is identical. The demo saves you five minutes of setup and costs you control: files are written to our machine, the rate limit is shared, and there is no SLA. Anything confidential or production-facing should be self-hosted.
Yes — edit-doc in append mode patches document.xml directly, so the original fonts, headers, footers and embedded images are untouched. Replace mode rebuilds the body and does not preserve formatting, so choose deliberately.
Through the remote-read bridge. CogniRunner mints a single-use URL and bearer for the attachment and hands them to the model; read-doc fetches the JSON envelope, decodes it into a per-call temp directory, parses it, and deletes the temp directory afterwards. Nothing is copied to a shared disk.
Yes. It is a standard MCP server over stdio, so LM Studio, Cursor, Cline, Roo Code, Continue, Qwen Code and any MCP host work with the same command-plus-args config.
It returns empty text rather than an error, because there is no text layer to extract. Check the result rather than assuming a read succeeded — and set a vision key if you need OCR.
Clone it, run it, fork it, ship it. The hosted demo is one deployment of the same repository — there is no paid tier holding anything back.