Notes from the work

One email when a tutorial or migration write-up goes live. Nothing else, and one click to leave.

LeanZero

Two people in Romania doing Atlassian migrations, Forge apps and practical AI work for teams that would rather talk to the person doing the job. Most of what we learn ends up on this site.

Services

  • Atlassian Migrations
  • Atlassian FastShift
  • Atlassian Maintenance
  • Forge App Development
  • AI Development Consultation

Topics

  • Jira
  • Jira Service Management
  • Confluence
  • Bitbucket
  • Atlassian Forge
  • Cloud Migration
  • Local AI
  • AI Coding
  • Certifications
  • Atlassian Team
  • All topics

Company

  • Blog
  • Tutorials
  • Contact

Community

  • Join Discord
  • Support this site

© 2026 LeanZero. All rights reserved.

Privacy PolicyTerms of Service24/7 SupportTrust CenterSecurityDPALegal notice
  1. Home
  2. Portfolio
  3. Mcp Doc Processor
MCP Server, Open Source, MIT

MCP Doc Processor

Set it up from scratch in five minutes, then read, create and edit PDF, DOCX, Excel and PowerPoint from any AI client.

Set it upView on GitHub
37-section manual 17 tools Node · JavaScript

Set it up in five minutes

Node 20 or newer, git, and about 750 MB of disk. No account, no API key, and nothing listening on the network.

1

Clone and install

terminal
git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git
cd leanzero-mcp-doc-processor
npm install
pwd    # copy this — the client config needs the absolute path

There is no build step — the server runs from source. Puppeteer downloads Chromium during install, which is the slow part and what renders create-pdf output.

2

Point your client at it

Claude Code
claude mcp add doc-processor --scope user \
  -- node /ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js
everything else — mcp.json
{
  "mcpServers": {
    "doc-processor": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {
        "DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents"
      }
    }
  }
}
3

Prove it works — before blaming your client

terminal
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node src/index.js 2>/dev/null | tail -1

# A healthy server answers with all 17 tools.
MCP over stdio is line-delimited JSON-RPC, so you can drive the server from a shell with no client involved. If this prints 17 tools, the install is sound and every remaining problem is client configuration. It is the single most useful debugging trick with this server.

Full detail — prerequisites, where files land, OCR, every client, and what each failure means — is in part 1 of the manual.

Which way should I run it?

Three deployments, one codebase. Most people want the first, and the manual documents all three end to end.

Local, over stdio

Recommended for almost everyone

Your MCP client launches the server as a child process. No port, no listener, no auth to manage — and documents land on your own disk, in your own project.

  • Two lines of JSON
  • Nothing on the network
  • Your files stay yours
Install it

An always-on HTTPS service

When something remote has to reach it

The HTTP transport under launchd or systemd, published through a Tailscale Funnel with per-tenant bearer tokens. This is the exact recipe our own Mac Studio runs.

  • launchd / systemd files
  • Tailscale Funnel, no open ports
  • Token minting and revocation
Run the server

The hosted demo

To evaluate it in a minute

One instance of this exact repo on our Mac Studio. Grab a free key and point any MCP client at it. An evaluation surface — not somewhere to put anything confidential.

  • Free key by email
  • No install at all
  • Not for production
Get a demo key

Claude Code

One `claude mcp add` command — stdio for a local build, or `--transport http` for the hosted endpoint.

Claude Desktop

One entry in `claude_desktop_config.json`. Use the absolute `node` path — a GUI app does not inherit your shell PATH.

LM Studio

The same `mcp.json` shape. This is also the bridge Forge apps reach through, so it is worth getting exactly right.

Cursor · Cline · Roo · Qwen

Any MCP client takes `command` + `args` + `env`. That is the whole contract.

Atlassian Forge apps

CogniRunner brokers every call from inside Forge, so your AI provider never sees the URL or the token.

The manual

Every question this server raises, answered in order: install it, connect your client, run it as a service the way we do, use all 17 tools, read the code, wire it into Jira through CogniRunner, and fix it when it misbehaves. 37 sections.

Written against the server's own source and against the live deployment behind the hosted demo — the launchd, Tailscale and hardening steps are the configuration actually running, with the secrets removed.

In this manual
  • What do I need before I start?
  • How do I install it?
  • How do I prove the install actually works?
  • Where do the files it creates actually land?
  • How do I turn on OCR for scanned PDFs?
  • How do I add it to Claude Code?
  • How do I add it to Claude Desktop?
  • How do I add it to LM Studio?
  • How do I add it to Cursor, Cline, Roo Code, or Qwen Code?
  • The server does not show up. What now?
  • Do I want stdio or HTTP?
  • How do I start the HTTP transport?
  • How do I keep it running on macOS? (launchd)
  • How do I keep it running on Linux? (systemd)
  • How do I publish it over HTTPS without opening a port?
  • How do I issue and revoke access tokens?
  • What must I check before this faces the internet?
  • Which tool do I call?
  • How do I read a document?
  • How do I create a document that does not look AI-generated?
  • How do I edit a document without wrecking it?
  • How do I make every document look like ours?
  • How do I fact-check a document against the live web?
  • What is in the repo, file by file?
  • What actually happens on one tool call?
  • How do the read and upload bridges work?
  • How do I add my own tool?
  • How do I know I have not broken it?
  • What does CogniRunner do with this MCP?
  • Which connection path do I want?
  • How do I connect it over the hosted bridge?
  • How do I connect it locally through LM Studio?
  • How does a rule read a Jira attachment?
  • What rules are actually worth building?
  • What is the hosted demo, and what is it not?
  • How do I connect to the hosted demo?
  • Troubleshooting

01What do I need before I start?

Node 20 or newer, git, about 750 MB of disk, and five minutes. No API key, no account, no LeanZero anything — the server is MIT-licensed and runs entirely on your machine.

MCP Doc Processor is a Node.js program that speaks the Model Context Protocol over stdin/stdout. Your AI client (Claude Code, Claude Desktop, LM Studio, Cursor, Cline, Qwen Code — anything that reads an mcp.json) launches it as a child process and talks JSON-RPC to it. That is the whole deployment model for the local case: there is no daemon to install, no port to open, and nothing listening on the network.

Requirements

WhatVersionWhyCheck it
Node.js20 or newer (we run 24 in production)package.json declares engines: { node: ">=20" }. The code uses ESM, top-level await and the built-in fetch.node --version
npm10 or newer (ships with Node 20+)Installs the dependency tree.npm --version
gitanyCloning the repo. A downloaded ZIP works too.git --version
Disk~750 MBMeasured on a clean clone: 185 MB of node_modules, plus ~565 MB of Chrome and chrome-headless-shell that Puppeteer downloads into ~/.cache/puppeteer for create-pdf. The browser cache is shared between clones, so a second checkout costs only the 185 MB.df -h .
An MCP clientanyClaude Code, Claude Desktop, LM Studio, Cursor, Cline, Qwen Code, Roo Code, or your own MCP host.—
NoteNothing here is optional-but-really-required. No API key is needed for 15 of the 17 tools. Two do need credentials: read-doc only when OCR-ing a scanned (image-only) PDF — see "How do I turn on OCR" below — and fact-check, which always needs a bearer for the web-search MCP plus a Serper key. Text PDFs, DOCX, XLSX and PPTX are parsed locally with no network access at all.

Windows, macOS and Linux all work. The examples below use macOS/Linux paths; on Windows use the same commands in PowerShell and write paths as C:\\Users\\you\\... with doubled backslashes inside JSON.

02How do I install it?

Clone, install, note the path — then prove it works before you touch a client config.

The five minutes

  1. 1Clone the repo. git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git then cd leanzero-mcp-doc-processor. The repository is public, so no GitHub account or token is needed.
  2. 2Install the dependencies. npm install. This pulls @modelcontextprotocol/sdk, docx, xlsx-js-style, pdf-parse, mammoth, pptxgenjs, marked, jszip and puppeteer. Puppeteer downloads Chrome into ~/.cache/puppeteer on install — that is the slow part and roughly 565 MB of the total, and it is what renders create-pdf output.
  3. 3Note the absolute path. Run pwd (macOS/Linux) or cd (Windows) and copy the result. Every client config below needs the absolute path to src/index.js; a relative path will not work because the client launches the process from its own working directory.
  4. 4Prove the server starts. Run the handshake in the next section. Do not skip it — a client that silently fails to launch the server looks exactly like a client that has no tools.

The whole install

git clone https://github.com/leanzero-srl/leanzero-mcp-doc-processor.git
cd leanzero-mcp-doc-processor
npm install
pwd    # copy this — you need the absolute path below
TipOn a slow link, or in CI, skip the Chromium download with PUPPETEER_SKIP_DOWNLOAD=1 npm install. Everything works except create-pdf, which will error when called. Add Chromium later with npx puppeteer browsers install chrome.

There is no build step. The server is plain ESM JavaScript and runs from source — src/index.js is the entry point for the stdio transport, src/server.js for the HTTP one. Nothing is transpiled or bundled, which means you can read and edit any file in the repo and restart your client to pick it up.

03How do I prove the install actually works?

Send a real MCP handshake down the pipe and count the tools. If you get 17 back, the server is genuinely working — before any client is involved.

MCP over stdio is line-delimited JSON-RPC. That means you can drive the server from a shell with echo and a pipe, with no client in the loop. This is the single most useful debugging trick with this server, and it is worth doing once now so you know what a healthy response looks like.

Handshake and list the tools

printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node src/index.js 2>/dev/null \
  | tail -1 | python3 -c 'import json,sys; print(len(json.load(sys.stdin)["result"]["tools"]), "tools")'
TipExpected output: 17 tools. If you get that, the server is installed correctly and every problem from here on is a client-configuration problem, not an install problem.

What each failure means

  • `Cannot find module` — you ran the probe from the wrong directory, or npm install did not finish. cd into the repo root and re-run npm install.
  • `SyntaxError: Unexpected token` / `await is only valid` — your Node is older than 20. Check with node --version and upgrade.
  • Nothing at all, hangs forever — the server is waiting for more input, which is correct behaviour. The printf above closes the pipe; if you typed the JSON by hand, press Ctrl-D.
  • Output is empty but exit code is 0 — you redirected the wrong stream. The protocol is on stdout, logs are on stderr; 2>/dev/null above is what keeps them apart.
CarefulNever add a `console.log` to `src/`. Anything written to stdout that is not JSON-RPC corrupts the protocol frame and the client drops the connection with a parse error. The repo ships a check for it — npm run lint:no-console-log exits non-zero and names the file — but nothing runs it for you: there is no CI configured in the repository, so run it yourself before committing. Use the logger in src/utils/logger.js, which writes to stderr.

04Where do the files it creates actually land?

By default, in the working directory the client launched the server from — which for a coding agent is your open project. DOC_OUTPUT_DIR pins it somewhere fixed instead.

This is the question people get wrong most often, so it is worth being precise. src/tools/utils.js resolves the output root in three steps, first match wins:

  1. 1A per-request override (X-Output-Dir header) — only exists on the HTTP transport, and it is sandboxed under a server-side base directory. A remote tenant can organise files into subfolders but cannot write anywhere it likes on the host.
  2. 2The `DOC_OUTPUT_DIR` environment variable — your own machine, your own choice, not sandboxed. Set this in the env block of your client config.
  3. 3`process.cwd()` — the default. For a stdio server launched by a coding agent, that is the workspace you opened the agent in, so generated documents appear right next to the code you are working on.
CarefulThe server also drops a `logs/` directory into its launch directory. Harmless, but it means a logs/ folder appears in whichever project your agent opened — observed on a clean install. Add it to that project's .gitignore, or set DOC_OUTPUT_DIR and launch from elsewhere.
CarefulA remote server writes to the remote host's disk. That is a client/server boundary, not a setting — X-Output-Dir on the hosted demo organises files on our Mac Studio, and you retrieve them over the signed /files/download link the tool returns. If you want documents to appear in your own Finder, run the server locally over stdio. There is no configuration that makes a remote process write to your laptop.

Inside the output root, the server enforces a docs/ folder and sorts documents into one of six category subfolders (contracts, technical, business, legal, meeting, research) chosen by the auto-categoriser in src/utils/categorizer.js. Pass enforceDocsFolder: false on a create call to opt out and write exactly where you asked.

Pin the output directory

{
  "mcpServers": {
    "doc-processor": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {
        "DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents"
      }
    }
  }
}

05How do I turn on OCR for scanned PDFs?

One environment variable, and you choose the provider: a cloud vision model, or a local one in LM Studio so no page image ever leaves your machine.

A PDF that was produced by a word processor contains a text layer, and pdf-parse reads it with no AI and no network. A PDF that was produced by a scanner is a bag of images: there is nothing to extract, and the only way to read it is to look at it. That is what the vision path in src/services/vision-service.js is for, and it is the only part of this server that ever calls out to a model.

Vision environment variables

VariableDefaultWhat it does
Z_AI_API_KEY(none)The key. ZAI_API_KEY and ANTHROPIC_AUTH_TOKEN are also accepted as fallbacks. Without one, OCR is simply off and everything else still works.
Z_AI_MODE(unset)ZAI (also Z_AI or Z; PLATFORM_MODE is read as a fallback) selects the Z.AI coding-plan endpoint https://api.z.ai/api/coding/paas/v4/ — not the standard one. Set Z_AI_CODING_PLAN=false alongside it for https://api.z.ai/api/paas/v4/. Anything else falls back to BigModel (open.bigmodel.cn).
Z_AI_BASE_URLauto from Z_AI_MODEOverrides the base URL outright. This is the local-OCR switch — point it at any OpenAI-compatible /v1/ endpoint.
Z_AI_CODING_PLAN(unset)Only read when Z_AI_MODE selects Z.AI. Anything other than the literal false keeps the coding-plan endpoint; set it to false if your key is for the standard API.
Z_AI_VISION_MODELglm-4.6vModel name sent in the request body.
Z_AI_TIMEOUT300000Request timeout in milliseconds. Vision on a long scan is slow; 5 minutes is deliberate.
SKIP_TABLE_EXTRACTIONtrueSet false to also run table reconstruction on page images. Slower, and worth it on financial scans.

Run OCR locally, with no cloud key at all

The vision client posts an OpenAI-shaped chat/completions request to Z_AI_BASE_URL. Any server that speaks that shape works — including LM Studio on your own machine. Load a vision-capable model in LM Studio, start its local server, and point the MCP at it. Page images then never leave the machine.

Local vision via LM Studio

{
  "mcpServers": {
    "doc-processor": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {
        "Z_AI_BASE_URL": "http://localhost:1234/v1/",
        "Z_AI_VISION_MODEL": "<the model id LM Studio shows>",
        "Z_AI_API_KEY": "lm-studio"
      }
    }
  }
}
CarefulThe trailing slash on `Z_AI_BASE_URL` is load-bearing. The client builds the URL as baseUrl + "chat/completions" with no separator, so http://localhost:1234/v1 (no slash) produces .../v1chat/completions and 404s.
CarefulPlaceholder keys are rejected on purpose. isConfigured() refuses any key containing the substring "api" or "key", so a literal your-api-key reads as not-configured rather than failing later with a confusing 401. If OCR says "not configured" while you are sure the variable is set, check the value for those substrings — lm-studio and a real token both pass.

To be precise about what leaves the machine: OCR never touches the web-search MCP, but it is not automatically local either — its one outbound call goes to Z_AI_BASE_URL, which is a cloud endpoint unless you point it at LM Studio as shown above. The other outbound path is fact-check, which calls the sibling web-search MCP — see the tools part.

06How do I add it to Claude Code?

One command. Pick the scope deliberately — that choice decides whether the server follows you between projects or stays with this one.

Add it

claude mcp add doc-processor --scope user -- node /ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js

Everything after -- is the command Claude Code will run. The -- is required: without it the CLI tries to parse node as its own argument.

Which scope?

--scopeWritten toUse it when
useryour home configYou want the tools in every project you open. This is the right default for a document server.
project.mcp.json in the repo, committedYour team should get the same server. Everyone still needs the repo cloned at the path you commit, so prefer npx-style or a path everyone shares.
localthis project only, not committedYou are trying it out. The default if you pass no scope.

With environment variables

claude mcp add doc-processor --scope user \
  -e DOC_OUTPUT_DIR=/Users/you/Documents/ai-documents \
  -e Z_AI_API_KEY=sk-... \
  -- node /ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js

Verify it, do not assume it

  • Run claude mcp list — the server should be listed and marked connected.
  • In a session, ask for the tool list, or just say "read this PDF" with a real path. A server that registered but crashes on launch shows up as connected-then-gone.
  • claude mcp remove doc-processor undoes it; re-add with a corrected path.
TipCommitting a `project` scope entry? Put the repo somewhere every teammate has it, or vendor the server as a dependency. A .mcp.json pointing at /Users/alice/... is a broken config for everyone except Alice.

07How do I add it to Claude Desktop?

Edit one JSON file and restart the app. The file location differs per OS, and "restart" means quit, not close the window.

Where the config lives

OSPath
macOS~/Library/Application Support/Claude/claude_desktop_config.json
Windows%APPDATA%\\Claude\\claude_desktop_config.json
Linux~/.config/Claude/claude_desktop_config.json

claude_desktop_config.json

{
  "mcpServers": {
    "doc-processor": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {
        "DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents"
      }
    }
  }
}
  1. 1Create the file if it does not exist. If it does, merge your entry into the existing mcpServers object rather than replacing the file — you will otherwise silently remove every other server you had.
  2. 2Use the absolute path to src/index.js. Claude Desktop launches from its own working directory, so ./src/index.js resolves to somewhere inside the app bundle.
  3. 3Fully quit Claude Desktop (Cmd-Q on macOS, right-click → Quit in the tray on Windows) and reopen it. Closing the window does not reload the config.
  4. 4Open the MCP/tools indicator in the composer. The doc-processor tools should be listed.
Careful`command: "node"` relies on `node` being on the app's PATH, which for a GUI app on macOS is not your shell's PATH. If the server does not appear, replace "node" with the absolute binary path — which node gives it, typically /usr/local/bin/node or /opt/homebrew/bin/node. This is the single most common Claude Desktop failure and it produces no visible error.

08How do I add it to LM Studio?

LM Studio reads the same mcp.json shape. This is also the path CogniRunner and any other Forge app depend on, so get it right once and it serves both.

  1. 1In LM Studio, open Program → Edit mcp.json (in older builds: the plugin/integrations panel).
  2. 2Add the doc-processor entry below, with the absolute path to src/index.js.
  3. 3Save. LM Studio reloads MCP servers on save — no restart needed.
  4. 4Load a model that supports tool calling, then ask it to read a file. If the model has no tool support, the server is connected but nothing will ever call it.

LM Studio mcp.json

{
  "mcpServers": {
    "doc-processor": {
      "command": "/usr/local/bin/node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {
        "DOC_OUTPUT_DIR": "/Users/you/Documents/ai-documents",
        "MCP_CLIENT_TYPE": "interactive"
      }
    }
  }
}

MCP_CLIENT_TYPE=interactive is worth setting here. It makes the create tools answer with a single line (Created: <path>) instead of the full metadata envelope. A human reading the model's reply wants the former; an agent parsing it wants the latter. See the clientHint section in the tools part for the per-call override.

TipUse the absolute node path in LM Studio. Same reason as Claude Desktop: a GUI app does not inherit your shell PATH. which node on this machine returns /usr/local/bin/node, which is what our own LM Studio config uses.
CarefulSetting this up so CogniRunner can use it? Name the entry `doc-reader`, not `doc-processor`. The name above is fine for your own LM Studio use, but CogniRunner looks the integration up by the exact key doc-reader — any other name and the tools are never offered to the model, with no error anywhere. See "How do I connect it locally through LM Studio?" in the CogniRunner part.

09How do I add it to Cursor, Cline, Roo Code, or Qwen Code?

They all read the same mcpServers object. Only the file path differs, and none of them need anything the others do not.

ClientWhere the config lives
Cursor.cursor/mcp.json in the project, or the global one under Settings → MCP
Cline (VS Code)cline_mcp_settings.json — open it from the extension's MCP Servers panel
Roo CodeSame shape, from the extension's MCP panel
Qwen Code.qwen/settings.json, under an mcpServers key
Continue~/.continue/config.json, under experimental.modelContextProtocolServers
Anything elseIf it speaks MCP over stdio, it takes command + args + env. That is the entire contract.

The block, unchanged, for all of them

{
  "mcpServers": {
    "doc-processor": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "env": {}
    }
  }
}
NoteTimeouts. A big scanned PDF with OCR can take minutes. Clients that expose a per-server timeout (Cline, Qwen Code, Roo Code) default to 60 s, which is too short. Set "timeout": 300000 alongside command if your client supports it.

10The server does not show up. What now?

Five causes account for almost every case, and each has a test that distinguishes it from the others in one command.

Work down this list — do not skip step 1

  1. 1Does the server start at all? Run the handshake probe from "How do I prove the install actually works?". If that fails, the problem is the install, not the client, and no amount of config editing will help.
  2. 2Is `node` findable by the client? Replace "command": "node" with the absolute path from which node. GUI apps do not inherit your shell PATH. This is the top cause on macOS.
  3. 3Is the path to `index.js` absolute and correct? node /the/exact/path/from/your/config should hang (waiting for input), not print Cannot find module. Ctrl-C to exit.
  4. 4Is the JSON valid? A trailing comma silently disables the whole config in most clients. python3 -m json.tool < your-config.json will tell you in one line.
  5. 5Did you actually restart? Claude Desktop and Cursor need a full quit. LM Studio reloads on save. Claude Code re-reads on the next session.

Read the server's own logs

# The server writes logs/server.log into the directory it was LAUNCHED from
# (process.cwd()) — for a stdio server that is your client's working directory,
# NOT the repo. For a coding agent it is the project you opened.
tail -f /THE/DIRECTORY/YOUR/CLIENT/LAUNCHED/FROM/logs/server.log

# Claude Desktop keeps per-server logs of its own:
tail -f ~/Library/Logs/Claude/mcp-server-doc-processor.log   # macOS
Careful"Unexpected token in JSON" from the client means something wrote non-protocol output to stdout. If you have edited src/, that is almost certainly a stray console.log. Run npm run lint:no-console-log — it exits non-zero and names the file.

11Do I want stdio or HTTP?

stdio if the AI runs on the same machine as your files. HTTP if something remote — a Forge app, a phone, claude.ai — needs to reach it. Most people want stdio and stop reading here.

stdio (src/index.js)HTTP (src/server.js)
Who launches itYour MCP client, as a child processlaunchd / systemd, as a long-running service
Network exposureNone. No port, no listener.A listener, and whatever you put in front of it
AuthNone needed — the process runs as youPer-tenant argon2 bearer tokens, or OAuth 2.1
Where files landYour disk, in the client's working directoryThe server's disk; you fetch them over a signed link
Who can use itOne machineAnything that can reach the URL
Set-up effortTwo lines of JSONThis whole part
TipRun both. They are separate entry points over the same tool registry, so a stdio config in your editor and an HTTP service for your Forge app can coexist against one clone of the repo with no conflict. That is exactly how this machine is set up.

The rest of this part is the real, working recipe from the Mac Studio that serves the hosted demo on this site: a launchd job, secrets in a mode-0600 env file outside the repo, and one Tailscale Funnel publishing both MCP servers on path prefixes. Nothing here is aspirational — it is the running configuration, with the secrets removed.

12How do I start the HTTP transport?

node src/server.js, bound to localhost, with a handful of environment variables. Confirm it with /healthz before you put anything in front of it.

Start it by hand first

cd /path/to/leanzero-mcp-doc-processor
PORT=10000 \
BIND_HOST=127.0.0.1 \
DATA_DIR="$HOME/Library/Application Support/mcp-doc-processor" \
DOC_PROCESSOR_ADMIN_TOKEN="$(openssl rand -base64 32)" \
node src/server.js

# In another terminal:
curl -s localhost:10000/healthz     # -> {"ok":true,...}

Server environment variables

VariableExampleWhat it does
PORT10000Listen port.
BIND_HOST127.0.0.1Bind to loopback. The proxy in front reaches it locally; nothing else should. Binding 0.0.0.0 puts the raw server on your LAN.
DATA_DIR~/Library/Application Support/mcp-doc-processorWhere tenants.json, the OAuth stores and the client output tree live. Keep it outside the repo so a git clean cannot delete your tokens.
PUBLIC_HOSThost.tailXXXX.ts.net:10000DNS-rebinding protection: the Host header must match. Include the port if the proxy forwards a non-default one.
ALLOWED_HOSTShost.tailXXXX.ts.net,host.tailXXXX.ts.net:443Extra accepted Host values, comma-separated. Needed when one server answers on both a path prefix (:443) and a dedicated port.
DOC_PROCESSOR_ADMIN_TOKENopenssl rand -base64 32Guards /v1/admin/*, the token-minting API.
TENANT_RATE_LIMIT60Requests per minute per tenant.
ISSUER_URL(unset)Set to the public https://… base to switch on the OAuth 2.1 authorization server for claude.ai web. Leave unset for bearer-only.
CLIENT_OUTPUT_BASE$DATA_DIR/client-outputRoot under which the sandboxed X-Output-Dir header may create subfolders.
PROVISION_SECRET(unset)Shared secret for /v1/provision, the self-service minting door. Leave unset unless you are running a signup flow of your own.
FILE_TOKEN_SECRETfalls back to PROVISION_SECRET, then DOC_PROCESSOR_ADMIN_TOKEN, then a random per-process secretHMAC key for the signed /files/download links. Set it explicitly. With none of the three set the server signs with a secret regenerated on every start, so every live download link 403s the moment launchd or systemd restarts the service — and KeepAlive/Restart=always make restarts routine.

Endpoints it exposes

RouteMethodPurpose
/healthzGETLiveness. Unauthenticated. The first thing to curl, always.
/mcpPOSTThe MCP endpoint itself. Bearer or OAuth.
/files/downloadGETSigned, expiring link to a file the tools produced. HMAC-SHA256, keyed on FILE_TOKEN_SECRET — falling back to PROVISION_SECRET, then DOC_PROCESSOR_ADMIN_TOKEN, then a random per-process secret (see the variable table). The link is built from PUBLIC_HOST, so it carries whatever port that variable names — see the note below.
/v1/admin/tenantsGET / POST / DELETEMint, list, rotate and revoke bearer tokens. Localhost-only by default.
/v1/provisionPOSTSelf-service minting, reachable through the funnel but guarded by a shared PROVISION_SECRET that only your own backend knows. This is what the key form on this page calls.
/.well-known/oauth-authorization-serverGETOAuth discovery — present only when ISSUER_URL is set.
CarefulThe download link inherits `PUBLIC_HOST`, port and all. Our instance sets PUBLIC_HOST=…ts.net:10000, so a create call returns https://…ts.net:10000/files/download?token=…. A browser fetches that happily (verified: HTTP 200, text/markdown) — but Forge egress cannot, because it only reaches port 443. If a CogniRunner rule should hand a user a working link, set PUBLIC_HOST to the bare :443 path-prefixed hostname rather than the dedicated port.

Three request headers are accepted alongside Authorization: X-ZAI-Key (the caller's own vision key, used for that request only and never stored), X-Output-Dir (a sandboxed subfolder name under CLIENT_OUTPUT_BASE), and the standard Mcp-Session-Id / Mcp-Protocol-Version. Each also has a query-string form (?zai_key=, ?output_dir=) for clients that cannot set headers.

13How do I keep it running on macOS? (launchd)

A LaunchAgent that runs a wrapper script, which sources a mode-0600 env file and execs node. Three files — and the reason for three rather than one is that secrets must not live in the plist.

The repo ships a plist template at deploy/com.leanzero.docprocessor.plist that puts the environment inline. That works, but it means secrets sit in a file launchd reads and launchctl print can echo. The setup actually running on our Mac Studio splits it in three, and that is the version worth copying.

The three files

  1. 1`~/Library/Application Support/mcp-doc-processor/server.env` — every variable, KEY=value, one per line. chmod 600. This is the only file that holds secrets, and it lives outside the repo so it can never be committed.
  2. 2`~/Library/Application Support/mcp-doc-processor/run-http-server.sh` — sources that env file with set -a, then execs node. chmod +x. The exec matters: it keeps node as the job's own PID so launchd's KeepAlive supervises the real process rather than a shell that has already exited.
  3. 3`~/Library/LaunchAgents/com.user.mcp-doc-processor.plist` — points at the wrapper, sets RunAtLoad and KeepAlive, and sends both streams to one log file. No secrets in it at all.

run-http-server.sh

#!/bin/bash
# Sources server.env so secrets stay out of the launchd plist, then execs node.
# launchd's KeepAlive restarts this on crash; `exec` keeps node as the job's PID.
set -euo pipefail

APP_DIR="$HOME/Library/Application Support/mcp-doc-processor"
REPO="/Users/you/Projects/leanzero-mcp-doc-processor"

set -a
# shellcheck disable=SC1091
source "$APP_DIR/server.env"
set +a

exec /usr/local/bin/node "$REPO/src/server.js"

~/Library/LaunchAgents/com.user.mcp-doc-processor.plist

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
  "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>
    <string>com.user.mcp-doc-processor</string>

    <key>ProgramArguments</key>
    <array>
        <string>/bin/bash</string>
        <string>/Users/you/Library/Application Support/mcp-doc-processor/run-http-server.sh</string>
    </array>

    <key>WorkingDirectory</key>
    <string>/Users/you/Projects/leanzero-mcp-doc-processor</string>

    <key>RunAtLoad</key>
    <true/>
    <key>KeepAlive</key>
    <true/>

    <key>StandardOutPath</key>
    <string>/Users/you/Library/Logs/mcp-doc-processor.log</string>
    <key>StandardErrorPath</key>
    <string>/Users/you/Library/Logs/mcp-doc-processor.log</string>
</dict>
</plist>

server.env

PORT=10000
BIND_HOST=127.0.0.1
PUBLIC_HOST=your-host.tailXXXX.ts.net:10000
ALLOWED_HOSTS=your-host.tailXXXX.ts.net,your-host.tailXXXX.ts.net:443
DATA_DIR="/Users/you/Library/Application Support/mcp-doc-processor"
TENANT_RATE_LIMIT=60
Z_AI_MODE=ZAI
DOC_PROCESSOR_ADMIN_TOKEN=<openssl rand -base64 32>

Install, start, watch

# The directory does not exist yet — create it before writing the two files into it.
mkdir -p ~/Library/Application\ Support/mcp-doc-processor

chmod 600 ~/Library/Application\ Support/mcp-doc-processor/server.env
chmod +x  ~/Library/Application\ Support/mcp-doc-processor/run-http-server.sh

launchctl load -w ~/Library/LaunchAgents/com.user.mcp-doc-processor.plist
launchctl list | grep mcp-doc-processor      # PID, then 0 = running clean
curl -s localhost:10000/healthz

tail -f ~/Library/Logs/mcp-doc-processor.log
Careful`KeepAlive` plus a crash-on-boot is a restart loop. If launchctl list shows a changing PID and a non-zero exit status, read the log before reloading — usually a bad path in run-http-server.sh or a missing server.env. Unload with launchctl unload ~/Library/LaunchAgents/com.user.mcp-doc-processor.plist while you fix it.
NoteA LaunchAgent runs when you are logged in. For a headless Mac that reboots unattended, either enable automatic login, or move the job to a LaunchDaemon in /Library/LaunchDaemons (root-owned, chown root:wheel, chmod 644) and set UserName in the plist.

14How do I keep it running on Linux? (systemd)

The same three-file shape: a unit, an EnvironmentFile at mode 0600, and a hardened service section.

~/.config/systemd/user/mcp-doc-processor.service

[Unit]
Description=MCP Doc Processor (HTTP transport)
Wants=network-online.target
After=network-online.target

[Service]
Type=simple
WorkingDirectory=%h/Projects/leanzero-mcp-doc-processor
EnvironmentFile=%h/.config/mcp-doc-processor/server.env
ExecStart=/usr/bin/node %h/Projects/leanzero-mcp-doc-processor/src/server.js
Restart=always
RestartSec=3

# Hardening. ReadWritePaths must cover EVERY directory the server writes:
#   1. DATA_DIR            — tenants.json, OAuth stores, client output
#   2. <WorkingDirectory>/logs — the server writes logs/server.log into its cwd
#   3. DOC_OUTPUT_DIR      — set it explicitly in server.env, or created
#                            documents land in the repo's docs/ folder
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=%h/.local/share/mcp-doc-processor
ReadWritePaths=%h/Projects/leanzero-mcp-doc-processor/logs

[Install]
WantedBy=default.target

Install and verify

# Create EVERY directory the unit references, including the repo's logs/ —
# systemd refuses to start a unit whose ReadWritePaths target does not exist.
# And set a LINUX DATA_DIR in server.env: the macOS example above uses an
# "Application Support" path, and left unset DATA_DIR falls back to the working
# directory, putting tenants.json in the repo where a git clean can delete it.
#   DATA_DIR=/home/you/.local/share/mcp-doc-processor
#   DOC_OUTPUT_DIR=/home/you/.local/share/mcp-doc-processor
mkdir -p ~/.config/mcp-doc-processor \
         ~/.local/share/mcp-doc-processor \
         ~/Projects/leanzero-mcp-doc-processor/logs

# ...write server.env into ~/.config/mcp-doc-processor/ first, then:
chmod 600 ~/.config/mcp-doc-processor/server.env
systemctl --user daemon-reload
systemctl --user enable --now mcp-doc-processor
systemctl --user status mcp-doc-processor
journalctl --user -u mcp-doc-processor -f

curl -s localhost:10000/healthz
Tip`loginctl enable-linger $USER` — without it a user service stops when your last session ends, which on a server means it dies the moment you log out of SSH.
CarefulEvery path in `ReadWritePaths=` must already exist, or systemd refuses to start the unit at all — which is why the install block creates them before daemon-reload. DATA_DIR itself is created on demand by the server (verified), but systemd needs it present to build the mount namespace.
CarefulTwo ways this unit bites, both verified in a Fedora 41 / systemd 256 container. The server writes logs/server.log into its working directory, not into DATA_DIR — so the repo's logs/ needs its own ReadWritePaths line, which is why the unit above has two. (1) If that directory does not exist, systemd will not even start the unit: Failed to set up mount namespacing: …/logs: No such file or directory, status=226/NAMESPACE, restart loop. (2) If the directory exists but the ReadWritePaths line is missing, ProtectHome=read-only makes it read-only and node dies on the log stream with code: 'EROFS', status=1/FAILURE — also a restart loop, and /healthz never answers. Neither failure mentions logging, which is what makes them slow to diagnose.
CarefulSet `DOC_OUTPUT_DIR` explicitly in the systemd env file. Left unset, the output root is the working directory, so created documents land in the repo's docs/ folder — which ProtectHome=read-only also blocks. Point it at a path inside ReadWritePaths and the whole class of problem goes away.
CarefulIf tenants vanish on restart, DATA_DIR is outside ReadWritePaths (or inside the repo, where a git clean can take it).

15How do I publish it over HTTPS without opening a port?

Tailscale Funnel. It gives you a real certificate on a real hostname with no port forwarding, no dynamic DNS and no router configuration — and one funnel can carry both MCP servers on path prefixes.

The server binds to 127.0.0.1. Funnel terminates TLS at Tailscale's edge and proxies to that loopback port. Your firewall stays closed, your home IP stays unpublished, and the certificate is issued and renewed for you.

Set it up

  1. 1Install and join. brew install tailscale (or your distro's package), sudo tailscale up. Then tailscale status and note the MagicDNS name, e.g. your-host.tailXXXX.ts.net.
  2. 2Enable MagicDNS and HTTPS certificates in the Tailscale admin console under DNS. Funnel cannot issue a certificate without them.
  3. 3Grant the funnel attribute in Access Controls: "nodeAttrs": [{ "target": ["autogroup:member"], "attr": ["funnel"] }].
  4. 4Publish the port. tailscale funnel --bg --set-path=/docproc 10000. The service is now reachable at https://your-host.tailXXXX.ts.net/docproc/mcp.
  5. 5Set `PUBLIC_HOST` and `ALLOWED_HOSTS` to match and restart the service, or DNS-rebinding protection will reject every request with a Host mismatch.
  6. 6Verify from off-network — a phone on cellular, not the LAN. curl -s https://your-host.tailXXXX.ts.net/docproc/healthz.

One funnel, both MCP servers

# This is the live routing on our Mac Studio.
tailscale funnel --bg --set-path=/docproc   10000   # doc-processor
tailscale funnel --bg --set-path=/websearch 8443    # web-search
tailscale funnel --bg --set-path=/lmstudio  1234    # LM Studio, for the Forge bridge

tailscale funnel status
# https://your-host.tailXXXX.ts.net (Funnel on)
# |-- /docproc   proxy http://127.0.0.1:10000
# |-- /lmstudio  proxy http://127.0.0.1:1234
# |-- /websearch proxy http://127.0.0.1:8443
CarefulPath routing changes the `Host` header your server sees. On :443 the client sends the bare hostname; on a dedicated funnel port it sends host:port. That is precisely why our server.env lists both forms in ALLOWED_HOSTS. Symptom of getting it wrong: /healthz works locally, and every request through the funnel returns a rebinding rejection.
NoteDo not want it on the public internet? Use tailscale serve instead of tailscale funnel — identical syntax, but reachable only from devices on your own tailnet. Bearer auth still applies. This is the right choice unless you specifically need claude.ai web or a Forge app to reach it.

16How do I issue and revoke access tokens?

One bearer per consumer, minted over a localhost-only admin API, stored as an argon2 hash. Shown once — there is no way to read one back.

Mint, list, rotate, revoke

ADMIN=<DOC_PROCESSOR_ADMIN_TOKEN>

# Mint — the bearer is returned ONCE and never stored in plaintext.
curl -s -X POST localhost:10000/v1/admin/tenants \
  -H "Authorization: Bearer $ADMIN" -H "Content-Type: application/json" \
  -d '{"displayName":"alice-laptop"}'

curl -s            localhost:10000/v1/admin/tenants                 -H "Authorization: Bearer $ADMIN"
curl -s -X POST    localhost:10000/v1/admin/tenants/<id>/rotate     -H "Authorization: Bearer $ADMIN"
curl -s -X DELETE  localhost:10000/v1/admin/tenants/<id>            -H "Authorization: Bearer $ADMIN"

How this is designed to fail safe

  • The admin API refuses proxied requests. Anything arriving with an X-Forwarded-For header — i.e. through the funnel — gets a 404, even with a correct admin token. You mint tokens by being on the machine. ALLOW_ADMIN_OVER_FUNNEL=true lifts that, and you should not.
  • Tokens are argon2 hashes in `tenants.json`. A stolen copy of that file does not yield working bearers.
  • One token per consumer, not one shared token. Revoking a laptop then costs one DELETE instead of a rotation everybody has to chase.
  • Rate limiting is per tenant (TENANT_RATE_LIMIT), so one noisy client cannot exhaust the server for the others.
TipLost a bearer? There is no recovery — that is the point of hashing them. Rotate the tenant and hand out the new one.

17What must I check before this faces the internet?

Nine things. They take ten minutes together and every one of them has bitten somebody.

  1. 1`BIND_HOST=127.0.0.1`. Confirm with lsof -nP -iTCP -sTCP:LISTEN | grep 10000 — you want 127.0.0.1:10000, not *:10000.
  2. 2`PUBLIC_HOST` set, and matching what clients actually send. Test through the funnel, not from localhost.
  3. 3A strong `DOC_PROCESSOR_ADMIN_TOKEN`. openssl rand -base64 32. Not a password you have used elsewhere.
  4. 4`ALLOW_ADMIN_OVER_FUNNEL` unset. Verify: curl https://your-host…/docproc/v1/admin/tenants -H "Authorization: Bearer $ADMIN" must return 404.
  5. 5`chmod 600` on `server.env`, `tenants.json`, and the OAuth stores. ls -l and read the mode; the server writes them 0600 but a hand-edit or a restore from backup can widen them.
  6. 6`DATA_DIR` outside the repo. A git clean -xfd in the repo must not be able to delete your tenant database.
  7. 7`TENANT_RATE_LIMIT` set to something you can actually serve. 60/min per tenant is our number for a Mac Studio.
  8. 8Log rotation. KeepAlive plus an unrotated log file fills a disk eventually. newsyslog on macOS, logrotate on Linux.
  9. 9Decide about OCR keys. Leaving Z_AI_API_KEY unset makes the server keyless — each caller brings their own via X-ZAI-Key, used for that request and never persisted. That is the posture we run, and it means a compromise of the host leaks no vision credential.
TipThe server tells you what you forgot — read the first two log lines. On boot it emits PUBLIC_HOST not set — DNS rebinding protection disabled and ISSUER_URL not set — OAuth disabled (static bearer only) when those are missing. If you see the first one on a public instance, stop and fix it before anything else.
CarefulFunnel is the public internet. Not "a bit exposed" — genuinely public, indexed by scanners within hours. The wall between a stranger and your documents is exactly: argon2 bearer tokens, the per-tenant rate limit, and Host checking. If any of the three is misconfigured, there is no second layer.
LimitWhat this server does not do: virus-scan uploads, sandbox the documents it parses, or bound total disk usage. Parsing is done by pdf-parse, mammoth and xlsx in-process. Treat a public instance as something you would be content to lose, and keep it off a machine that holds anything else you care about.

18Which tool do I call?

Seventeen tools, four jobs: read something, make something, change something, or keep track of what you made. This table is the whole decision.

ToolModes / actionsCall it when
read-docsummary, indepth, focusedYou need what is inside a PDF, DOCX, XLSX or PPTX. Always before an edit.
detect-format—The user did not say which format they want. Returns a ready-to-use creation plan.
create-doc—An editable Word deliverable someone will keep working on.
create-pdf—A final deliverable to print, sign or send. Rendered through headless Chromium.
create-markdown—A README, spec, runbook or anything that lives in a repo.
create-excel—Tabular or numeric data, multiple sheets, formulas.
create-pptx—A real editable deck — one slide per ## heading, with speaker notes.
edit-docpreview, append, replace, styleChange a DOCX without destroying its formatting.
edit-excelpreview, append-rows, append-sheet, replace-sheetAdd to a workbook while keeping its styles.
edit-pptxpreview, append-slides, replace-slideChange an existing deck.
fact-check—Verify a document's claims against the live web, with citations.
list-documents—Find something you made before, or check for a duplicate.
list-templates—See which blueprints a create call can validate against.
dnainit, get, evolve, save-memory, delete-memoryMake every document come out looking like yours.
blueprintlearn, list, deleteCapture the structure of a document you like and enforce it later.
drift-monitorwatch, checkGet told when a document's structure changes underneath you.
get-lineage—Answer "where did this document come from?"
NoteOld tool names still work. get-doc-summary, get-doc-indepth, get-doc-focused, init-dna, get-dna, evolve-dna, learn-blueprint, list-blueprints, watch-document, check-drift and search-registry are all accepted as aliases, so a prompt written against an older version does not break. `save-memory` and `delete-memory` are not — they are dna actions, and calling either as a tool returns Unknown tool.

19How do I read a document?

read-doc with the mode that matches the question. Using indepth when you needed summary burns context for nothing; using summary before an edit loses the structure the edit depends on.

The three modes

ModeReturnsUse for
summaryOverview, page/sheet/slide count, a content preview"What is this file?" — the cheap default.
indepthFull text, structure, metadata, tablesAnything you are about to edit, or a real analysis. Expensive in context.
focusedAn answer to userQuery, sourced from the document"What is the payment term in this contract?" — one question, one answer.

The three shapes

{ "filePath": "/path/to/contract.pdf", "mode": "summary" }

{ "filePath": "/path/to/contract.pdf", "mode": "indepth" }

{ "filePath": "/path/to/contract.pdf",
  "mode": "focused",
  "userQuery": "What are the payment terms and the notice period?" }

What it can open

  • PDF — text layer via pdf-parse; image-only pages via the vision path when a key is configured. A scanned PDF with no vision key returns empty text, not an error, so check the result rather than assuming.
  • DOCX — mammoth for text, plus direct XML reading for structure, headers and footers.
  • XLSX / XLS — every sheet, with cell values and formulas.
  • PPTX — a per-slide transcript: titles, bullets, and speaker notes, plus the slide count.
TipAlways `indepth` before `edit-doc`. The edit tools patch XML in place; they need to know what is already there. Editing off a summary read is how you end up with a document that has two conclusions.

Reading a file that is not on this machine

read-doc also takes url + authHeader instead of filePath. The endpoint must return { data: <base64>, filename, mimeType, size } as JSON. The server decodes it into a per-call temp directory, runs the normal pipeline, and deletes the temp directory afterwards even if parsing throws. This is how CogniRunner hands a Jira attachment to the MCP without the file ever touching a shared disk — the wire contract is documented in the code part.

20How do I create a document that does not look AI-generated?

Write the body as markdown, pick a style preset, and let Document DNA supply the header, footer and house style. The difference between a usable deliverable and an obvious one is almost entirely here.

  1. 1Let `detect-format` choose the format if the user did not name one. It weighs the verbs ("print", "sign", "send to the client" → PDF; "editable", "draft", "template" → DOCX; "README", "spec", "for the repo" → Markdown; "budget", "tracker", "dataset" → Excel) and returns { format, suggestedTool, stylePreset, category, docType, confidence, reason } that you pass straight through.
  2. 2Write the body as markdown in `content`. Headings, bullets, numbered lists, blockquotes, links, code blocks and pipe tables are all parsed and rendered as native document features — real Word tables, real numbered lists. A wall of plain text renders as a wall of plain text.
  3. 3Give it a specific title. The rejected set is literal: Untitled, Untitled Document, New Document, Document, Doc, File, Output, Temp, Tmp, New File, Unnamed, No Title. Anything else passes — Report is accepted, so a vague-but-unlisted title still gets through. The title becomes the H1 and the registry key.
  4. 4Pick a `stylePreset` — or leave it and let DNA decide. The default is claude-like.
  5. 5Use `dryRun: true` first on anything long. It returns the full preview and the formatting-quality score without writing a file.

A real create-doc call

{
  "title": "Q3 Platform Migration Readiness Review",
  "stylePreset": "professional",
  "category": "technical",
  "tags": ["migration", "q3"],
  "content": "## Summary\n\nThe migration is **on track** for the 14 November window.\n\n## Open risks\n\n| Risk | Owner | Status |\n|---|---|---|\n| Attachment volume | Platform | Mitigated |\n| SSO cutover | Identity | Open |\n\n## Next steps\n\n1. Freeze schema changes on 1 November\n2. Dry-run the cutover on staging\n3. Confirm the rollback window with Ops",
  "header": { "text": "Acme Corp — Internal" },
  "footer": { "text": "Confidential", "pageNumbers": true }
}

The eight style presets

PresetFontBodyMade for
claude-likeCalibri11ptThe modern blue-accented default
professionalGaramond11ptExecutive summaries, formal reports
technicalArial11ptAPI docs, specs, manuals
legalTimes New Roman12ptContracts, agreements, briefs
businessCalibri11ptProposals, go-to-market plans
minimalArial11ptEveryday, clean documents
casualVerdana12ptInternal comms, team updates
colorfulArial12ptPresentations, marketing

Things worth knowing before you are surprised by them

  • Duplicate titles are refused by default. You get { duplicate: true, existingPath } back and should switch to edit-doc. Pass preventDuplicates: false to override.
  • Every response carries `formattingQuality` and sometimes `formatSuggestion`. If formatSuggestion is set, the content genuinely suits a different format — heed it rather than shipping a spreadsheet as a Word file.
  • `create-excel` takes formulas. [["Month","Units","Price","Revenue"],["Apr",120,9.99,"=B2*C2"]] produces a live formula, not a string.
  • `create-markdown` can emit a table of contents and YAML frontmatter. toc: true builds anchor links from the H2/H3s; frontmatter: { title, date, tags } feeds Hugo, Jekyll or Astro.
  • `create-pptx` builds one slide per `## ` heading, with bullets, native slide tables, native charts and speaker notes. It opens in PowerPoint, Keynote and Google Slides.
Note`detect-format` routes slide requests to `create-pptx`. format-router.js maps the PPTX format straight to that tool and returns a note explaining the choice, so the router can be trusted for decks. Reach for create-pdf instead only when the deck must be fixed-layout for printing or signing.
Note`clientHint` decides how chatty the answer is. "interactive" returns one line (Created: <path>); "agent" returns the full metadata envelope (enforcement, style config, lineage, memories applied) and is the default; "auto" guesses from the input shape and then falls back to MCP_CLIENT_TYPE. Set MCP_CLIENT_TYPE=interactive server-side when a human reads the output directly.

21How do I edit a document without wrecking it?

edit-doc in append mode patches the DOCX XML directly, so fonts, headers, footers and images survive. replace does not — know which one you asked for.

ToolActionWhat survives
edit-docpreviewNothing is written — it returns the document's current content, the same safety valve edit-pptx has.
edit-docappendEverything. Content is inserted by patching document.xml in place — original styling, headers, footers and embedded images are untouched.
edit-docreplaceThe file, not its formatting. The body is rebuilt from what you supply.
edit-docstyleThe content; the styling is re-applied. Undocumented elsewhere — read the handler before relying on it.
edit-excelpreviewNothing is written — it returns the workbook's current content.
edit-excelappend-rowsSheet styles, column widths, existing formulas.
edit-excelappend-sheetThe whole workbook; a new sheet is added.
edit-excelreplace-sheetThe workbook; one named sheet's data is replaced.
edit-pptxpreviewNothing is written — it returns the deck's current text and notes.
edit-pptxappend-slidesThe existing deck; new slides are added.
edit-pptxreplace-slideThe deck is rebuilt from its extracted text and notes with one slide changed.
Careful`edit-pptx` `replace-slide` rebuilds the deck from the text and speaker notes it can extract. Custom layouts, embedded video and hand-placed graphics do not survive that round-trip. append-slides does not rebuild and is safe. Preview first, always.

The safe edit sequence

  1. 1read-doc with mode: "indepth" — know exactly what is in there.
  2. 2edit-doc with action: "append" and the smallest change that does the job.
  3. 3read-doc again on the result if the edit mattered. The tools report success on a write; they do not assert that the document still reads correctly.

22How do I make every document look like ours?

Document DNA is a project-level file that supplies the company name, header, footer and style preset to every create call, so nobody has to remember them. It also learns.

  1. 1Initialise once per project. dna with action: "init" and your company name, preferred preset, header and footer defaults. It writes .document-dna.json in the output root.
  2. 2Then do nothing. Every create-doc / create-pdf / create-markdown / create-excel call reads DNA and fills in any field you did not pass explicitly. An explicit argument always wins.
  3. 3Save preferences as memories. dna save-memory with something like "Always use 1-inch margins for contracts". Memories persist in the DNA file and are surfaced to the agent as context.
  4. 4Let it evolve. dna evolve analyses what you have actually been producing and proposes mutations — "80% of your documents use the business preset". Pass apply: true to accept the top suggestion.

DNA inherits across three levels, each overriding the one before: hardcoded system defaults, then project DNA (.document-dna.json), then user DNA (.document-user.json). A field missing at one level falls through to the next, so a personal preference can override a project default without editing the shared file.

What DNA evolution also produces

ConceptStored inWhat it gives you
Blueprint.document-blueprints.jsonA structural template — required headings, expected sections — learned from a real document or auto-detected during evolve. Pass blueprint: "<name>" to a create call and the structure is validated.
Registrydocs/registry.jsonEvery document created, with category, tags, title and path. Backs list-documents and the duplicate check.
Lineagethe registryThe provenance chain: which sources informed a document, and what was derived from it. get-lineage walks it.
Driftfingerprintsdrift-monitor watch records a structural fingerprint; check tells you what changed since.
TipBlueprints are the answer to "why does every report come out differently?" Learn one from the report you liked (blueprint learn), then require it. Documents that miss a mandated section come back with a validation warning instead of quietly shipping.

23How do I fact-check a document against the live web?

fact-check extracts the document's claims, sends each one to the sibling web-search MCP, and reports what the live web says — with sources. It is the only tool whose core function is a call to another service.

This is a cross-MCP tool. src/services/web-search-client.js calls the MCP Web Search server per claim, gathers cited sources, and can write the whole verification out as a report document. It is the reason the two servers are documented as a pair rather than as two unrelated projects.

  1. 1Set up MCP Web Search as a reachable HTTP instance. fact-check connects over Streamable HTTP only — src/services/web-search-client.js constructs no stdio transport at all, so a local stdio build of web-search cannot be used here.
  2. 2Set all three of WEB_SEARCH_MCP_URL (its /mcp URL), WEB_SEARCH_BEARER (a tenant bearer for it) and SERPER_API_KEY (your own Serper key). The bearer and the Serper key are both mandatory — the tool returns webSearchBearer is required / serperKey is required and does nothing without them. Both can also be passed per call as webSearchBearer and serperKey. With WEB_SEARCH_MCP_URL unset the client falls back to LeanZero's hosted web-search endpoint.
  3. 3Call fact-check with the document path. Each extracted claim is searched, and the result carries the sources it found.
  4. 4Optionally have it write a verification report — it goes through the normal create pipeline, so it inherits your DNA and style.
CarefulFact-checking sends claim text to a search engine. That is an outbound disclosure of document content. On anything confidential, run the web-search MCP yourself rather than pointing at a shared instance, and read its own privacy notes first.
NoteNo web-search MCP configured? fact-check fails cleanly and every other tool carries on working. It is the only tool that cannot function at all without a second service — though read-doc also reaches out when given a url, or when OCR runs on a scanned PDF, and every create tool reaches out when given an uploadUrl.

24What is in the repo, file by file?

About 18,000 lines of plain ESM JavaScript in 59 files, with no build step. Two entry points share one tool registry; everything else is a leaf.

The tree

leanzero-mcp-doc-processor/
├─ src/
│  ├─ index.js            # ENTRY: stdio transport. ~35 lines.
│  ├─ server.js           # ENTRY: Express HTTP transport, auth, rate limit.
│  ├─ tool-registry.js    # The 17 schemas + the dispatch table. Start here.
│  ├─ auth.js             # argon2 tenant bearers, tenants.json
│  ├─ oauth/              # OAuth 2.1 (PKCE + dynamic client registration)
│  ├─ tools/              # One file per tool + shared helpers
│  ├─ services/           # Business logic: DNA, lineage, drift, OCR, routing
│  ├─ parsers/            # pdf / docx / excel / pptx readers
│  └─ utils/              # logger, registry, categoriser, request context
├─ test/                  # 14 suites, node:test
├─ deploy/                # launchd plist template + runbook
└─ docs/                  # Generated output + registry.json

The files that carry the weight

FileWhat lives there
src/tool-registry.jsRead this first. Every tool's JSON schema and description, plus the CallToolRequestSchema dispatch. The descriptions are long on purpose — they are the model's only instructions, so routing quality is decided in this file.
src/tools/styling.jsThe eight presets and every colour, spacing and table constant. Change the house look here.
src/tools/docx-patch.jsThe XML patching that makes edit-doc append non-destructive. The most delicate file in the repo.
src/utils/markdown-formatter.jsMarkdown → native document features. Why a pipe table becomes a real Word table.
src/utils/dna-manager.js + dna-inheritance.js + dna-schema.jsDocument DNA: read, merge across three levels, validate.
src/utils/registry.jsdocs/registry.json — the document index, duplicate detection and lineage store.
src/utils/request-context.jsAn AsyncLocalStorage carrying per-request values (zaiKey, outputDir) from the HTTP layer down into services without threading them through every signature.
src/services/format-router.jsdetect-format — the heuristics that pick DOCX vs PDF vs Markdown vs Excel.
src/services/vision-service.jsThe only outbound call in the read path. OpenAI-shaped chat/completions.
src/services/web-search-client.jsThe cross-MCP client fact-check uses to reach MCP Web Search.
Tip`src/index.js` is 35 lines and worth reading in full. It constructs the SDK Server, registers the tools, captures the connecting client's identity in oninitialized, and connects a StdioServerTransport. That is the entire stdio server — everything else is tools.

25What actually happens on one tool call?

Nine steps from JSON-RPC frame to file on disk. Knowing this sequence is what turns a confusing failure into an obvious one.

  1. 1Transport receives the frame. stdio reads a line from stdin; HTTP takes a POST /mcp body after mcpAuth (in src/server.js) has resolved a tenant — trying an OAuth access token first when ISSUER_URL is set, then falling back to requireBearer in src/auth.js, which verifies the bearer against the argon2 hash in tenants.json — and the per-tenant rate limiter has allowed it.
  2. 2Request context is established (HTTP only). X-ZAI-Key and X-Output-Dir are pulled off the request and stored in the AsyncLocalStorage in request-context.js, so nothing downstream needs them passed explicitly.
  3. 3The SDK routes CallToolRequestSchema to the handler registered in tool-registry.js.
  4. 4The client profile resolves. client-profile.js maps the clientInfo.name from the handshake to a cohort, which decides whether the response is chatty or terse when clientHint is auto.
  5. 5Arguments are validated against the tool's JSON schema by the SDK; a mismatch comes back as a protocol error, not a crash.
  6. 6Document DNA is merged in for create calls: system defaults, then .document-dna.json, then .document-user.json, with anything explicit in the call winning over all three.
  7. 7The output path is resolved through getOutputRoot() — the three-step precedence described in the setup part — then run through the docs/<category>/ enforcement and the duplicate check.
  8. 8The work happens. A parser reads, or docx / pptxgenjs / xlsx-js-style / headless Chromium writes.
  9. 9The registry is updated, lineage is recorded, and the response is shaped by clientHint. If uploadUrl was supplied, the file is POSTed to it before the response is returned.
NoteEverything logs to stderr, never stdout. src/utils/logger.js is the only sanctioned output path. This is not a style preference — stdout is the protocol channel on the stdio transport, and one stray console.log corrupts the frame and drops the connection. npm run lint:no-console-log enforces it.

26How do the read and upload bridges work?

Two mirror-image HTTPS contracts that let the server read from, and write to, any authenticated endpoint you control. This is the extension point that makes it useful inside someone else's platform.

doc-processor does not know what CogniRunner is. It knows one JSON envelope in each direction. Anything that speaks that envelope — a Forge web trigger, a Cloudflare Worker, a Lambda, an Express route — plugs in. CogniRunner is the reference implementation, not a special case.

Reading: url + authHeader on read-doc

The remote path fires only when both url and authHeader are present and filePath is not. Existing local callers are completely unaffected. Your endpoint returns HTTP 200 application/json with { data: "<base64>", filename, mimeType, size }. The server decodes into a unique per-call temp directory under os.tmpdir(), runs the normal extraction pipeline, and removes the temp directory in a finally — including when the pipeline throws.

Uploading: uploadUrl + uploadAuthHeader on any create tool

POST <uploadUrl>
Authorization: <uploadAuthHeader>
Content-Type: application/json

{ "data": "<base64 file bytes>",
  "filename": "q3-review.docx",
  "mimeType": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
  "size": 27834 }

--> 200 application/json
{ "success": true,
  "attachment": { "id": "...", "filename": "...", "size": 27834,
                  "mimeType": "...", "content": "<user-visible URL>" } }

Failure statuses, surfaced verbatim as uploadError

StatusMeans
400Malformed envelope
401Bearer mismatch
404Token expired or already consumed (single-use capability)
413Payload too large for the receiver
415Disallowed mimeType or extension
502The receiver's own upstream failed (e.g. Jira)
500Unexpected receiver error

Guarantees the sender makes — worth knowing before you build a receiver

  • HTTPS only. A non-https:// URL is rejected before fetch is called.
  • No redirects. redirect: "error". Your receiver must answer directly — a 3xx aborts the call.
  • No auto-retry on any 4xx or 5xx. Single-use capability semantics are honoured end to end, so a one-shot token really is one shot.
  • The auth header is never logged, at any level, and the URL's ?t= token is redacted — only host and path reach the log.
  • Nothing is cached. Bytes, URL and auth header live for the duration of one call.
  • Sizes are bounded. READ_DOC_MAX_BYTES (default 50 MB) on the way in, WRITE_DOC_MAX_BYTES (default 25 MB) on the way out. The write cap is half the read cap because Forge web-trigger payload limits are tighter.
  • Timeouts: 30 s on the remote read, 60 s on the upload.
  • The local file survives a failed upload, so your caller can retry or fall back to the path.
TipBuilding a receiver? Mint the URL and bearer per request, store them with a short TTL, delete on consume, and compare the bearer in constant time. Return 4xx/5xx with a JSON body — never a redirect.

27How do I add my own tool?

Six steps across four files, no framework. The hardest part is writing a description the model will actually route on correctly.

  1. 1Write the handler. A new file in src/tools/, exporting one async function that takes the validated arguments object and returns a result object. Import helpers from src/tools/utils.js — getOutputRoot(), enforceDocsFolder() — rather than reimplementing path handling.
  2. 2Import it in `src/tool-registry.js` alongside the others at the top.
  3. 3Add the schema to the tools array: name, a long description, and an inputSchema with explicit enums wherever a value is constrained.
  4. 4Add the dispatch case in the CallToolRequestSchema handler.
  5. 5Add a test under test/ using node:test, and wire it into test:all in package.json.
  6. 6Restart your MCP client. Servers are launched at client start; a running client will not see the new tool.
CarefulThe description is the model's only instruction manual. It never reads your source. Look at how create-doc and create-pdf describe themselves: each says what it is FOR, and explicitly what it is NOT for, with the other tool named. That negative half is what stops a model shipping a spreadsheet as a Word file — the routing quality of the whole server lives in those strings.

Conventions to keep

  • Never `console.log`. Use log() from src/utils/logger.js, and run npm run lint:no-console-log before you commit — there is no CI in the repo to run it for you.
  • Never write outside `getOutputRoot()`. It is what keeps a hosted tenant sandboxed.
  • Register what you create via src/utils/registry.js, or list-documents, duplicate detection and lineage will not see your files.
  • Return a result object, not a string. The registry shapes the response envelope according to clientHint; a raw string bypasses that.
  • Accept `clientHint` if the tool produces something a human might read directly.

28How do I know I have not broken it?

Fourteen suites, one command, no external services required. Run them before and after any change — they take under a minute.

Run everything

npm run test:all              # all 14 suites in sequence
npm run lint:no-console-log   # fails if any src/ file has a console.log

# Or one at a time, when you know what you touched:
npm run test:read-doc         # read-doc URL-fetch path
npm run test:schemas          # MCP schema invariants + detect-format E2E
npm run test:render           # markdown -> DOCX round-trip
npm run test:upload           # the upload bridge
npm run test:http             # the HTTP transport
npm run test:auth             # tenant bearer auth
npm run test:oauth            # the OAuth 2.1 flow
npm run test:pdf              # PDF parsing
npm run test:create-pdf       # PDF rendering (needs Chromium)
npm run test:create-pptx      # deck generation
npm run test:fact-check       # the cross-MCP fact-check path
npm run test:format-router    # detect-format heuristics

What each suite is actually protecting

SuiteThe regression it exists to catch
test:schemasA tool schema drifting from its handler — the failure mode where the model sends a valid-looking argument that is silently ignored.
test:renderMarkdown that stops becoming native document features: a table rendering as a paragraph of pipes.
test:uploadThe bridge contract. Redirect handling, retry suppression, size caps, and that the auth header never reaches a log.
test:authThat a wrong, absent or revoked bearer is refused — the wall in front of a public instance.
test:read-docThe remote-read path: temp-directory cleanup on the throwing case, and the size cap being enforced before materialising the file.
test:format-routerdetect-format routing. Every heuristic change risks sending decks to Excel.
Careful`test:create-pdf` needs Chromium. If you installed with PUPPETEER_SKIP_DOWNLOAD=1, that suite fails and it is not telling you anything about your change. Install the browser (npx puppeteer browsers install chrome) or skip that one suite knowingly.

29What does CogniRunner do with this MCP?

It turns a Jira workflow rule into something that can read the issue's attachments and produce a real deliverable back onto the issue. That is the whole point of wiring the two together.

CogniRunner is LeanZero's AI workflow app for Jira — validators, conditions and post-functions on transitions. On its own, a rule can reason about field text. With doc-processor connected, the same rule can open the PDF somebody attached, decide whether it says what it needs to say, and attach a generated report back to the issue.

What the integration exposes

ToggleTools the model may callWhat that enables
doc-reader (read)read-docThe rule can read the content of attached PDFs, DOCX and XLSX — not just their filenames.
doc-writer (write, sub-toggle, off by default)create-doc, create-pdf, create-markdown, create-excel, create-pptx, fact-check, list-templatesThe rule can produce a document and attach it to the issue. Requires the read toggle to be on as well.
NoteThe tool list is curated on purpose — 8 of the 17 are exposed, not 14. read-doc on the read toggle, plus create-doc, create-markdown, create-excel, create-pdf, create-pptx, fact-check and list-templates when doc-writer is on. The nine dropped are detect-format, edit-doc, edit-excel, edit-pptx, list-documents, dna, blueprint, drift-monitor and get-lineage — so a rule cannot edit an existing document, and the whole DNA/blueprint/lineage family is out of reach. None of them is usable in a single-shot workflow run, and a smaller surface is a better prompt.

The security story is the reason it is built this way: CogniRunner brokers every call from inside Atlassian Forge. Your AI provider receives the tool's result, never the MCP's URL, never the bearer token, and never the raw attachment.

30Which connection path do I want?

Two, and they are not interchangeable. The hosted-bridge path works with every AI provider; the local path works only with LM Studio but keeps the documents on your own machine.

Hosted bridgeLocal via LM Studio
How it connectsCogniRunner dials the MCP's HTTPS URL from ForgeLM Studio loads the MCP from your mcp.json over stdio
Works withEvery provider — Anthropic, OpenAI, Azure, OpenRouter, Bedrock, Forge LLMLM Studio only
What you configureService URL + Tenant Bearer, in the CogniRunner admin panelAn mcp.json entry on the LM Studio host
Where documents are writtenOn the MCP server's diskOn the LM Studio machine — yours
Who must run a serverYou, or LeanZero's hosted demoYou
CarefulDo not mix the two across enabled MCPs. LM Studio cannot combine native local plugins and hosted-bridge function tools in a single request. If you enable one MCP locally and another hosted, CogniRunner routes all of them through the hosted bridge — which means the "local" one then also needs a Service URL and Bearer, or it silently stops working. Keep every enabled MCP on the same side.

31How do I connect it over the hosted bridge?

Paste a URL and a bearer into the CogniRunner admin panel. The one thing that catches people is the port — Forge egress reaches 443 and nothing else.

  1. 1Get a URL and a bearer. Either take a free demo key from the "Try it without installing anything" section on this page, or self-host and mint your own tenant token (see "How do I issue and revoke access tokens?").
  2. 2Open CogniRunner's admin panel → Settings → MCP Integrations.
  3. 3Find the doc-processor card and turn the integration on.
  4. 4Paste the Service URL — it must start with https://. For the hosted demo that is https://worksmacstudio.tailfc4700.ts.net/docproc/mcp.
  5. 5Paste the Tenant Bearer. It must be at least 16 characters. Once saved it is masked and never returned to the UI — you can edit the URL later without re-entering it.
  6. 6Turn on the doc-writer sub-toggle if rules should also produce documents. It is off by default, and it needs the read toggle on.
  7. 7Use the card's probe to confirm CogniRunner can actually reach the server before you build a rule on it.
CarefulSelf-hosting? Your funnel must answer on port 443. The addresses CogniRunner may reach are fixed by the installed app: a *.ts.net Tailscale Funnel URL on port 443, or context7's own host. Forge egress does not reach 8443 or 10000, and an arbitrary URL like https://mycompany.com/mcp is blocked outright. That allow-list ships inside the app and cannot be changed without redeploying CogniRunner, which an installer cannot do.
TipThat is exactly why our own funnel uses path prefixes. tailscale funnel --bg --set-path=/docproc 10000 publishes the loopback service on https://<host>.ts.net/docproc — port 443 on the outside, port 10000 on the inside. Remember to add the bare hostname to ALLOWED_HOSTS as well as the host:port form, because the two routes send different Host headers.

32How do I connect it locally through LM Studio?

An ordinary mcp.json entry on the LM Studio host — with one non-obvious requirement: the entry has to be named doc-reader.

CarefulThe entry name must be exactly `doc-reader`. Not doc-processor, not docs. CogniRunner looks the integration up by that key, and any other name means the tools are simply never offered to the model — with no error anywhere.

LM Studio mcp.json — local stdio

{
  "mcpServers": {
    "doc-reader": {
      "command": "/usr/local/bin/node",
      "args": ["/ABSOLUTE/PATH/TO/leanzero-mcp-doc-processor/src/index.js"],
      "timeout": 300000,
      "env": {
        "MCP_CLIENT_TYPE": "interactive"
      }
    }
  }
}

LM Studio mcp.json — remote HTTP (LM Studio 0.3.17+)

{
  "mcpServers": {
    "doc-reader": {
      "url": "https://your-host.tailXXXX.ts.net/docproc/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_TENANT_BEARER"
      }
    }
  }
}

Then, on the CogniRunner side

  1. 1Set the provider to LM Studio and give CogniRunner your LM Studio's Tailscale Funnel URL (https://…ts.net) — it must be HTTPS, on *.ts.net, and never localhost.
  2. 2Turn on "Serve on Local Network" in LM Studio's developer settings, or the funnel relay cannot reach it.
  3. 3In the doc-processor card, switch on "Run locally via LM Studio (mcp.json)".
  4. 4Make sure every other enabled MCP is also set to local — see the mixing warning above.
  5. 5Load a tool-capable model and run the card's probe.

33How does a rule read a Jira attachment?

CogniRunner mints a one-shot URL and bearer per attachment and hands them to the model. Nothing is copied to a shared disk, and the capability dies on first use.

  1. 1The validator runs and sees the issue's attachments.
  2. 2For each one, CogniRunner mints a single-use capability: a URL with a short-lived token, plus a bearer, both bound to this issue and this run.
  3. 3Those values go into the model's prompt. The model calls read-doc with url + authHeader — exactly as given.
  4. 4doc-processor fetches the JSON envelope, decodes the base64 into a per-call temp directory, parses it, and deletes the temp directory afterwards.
  5. 5The extracted content goes back to the model. The capability is now spent.
CarefulNever retry a 404 on one of these URLs. The capability is single-use, so a 404 means it has already been consumed — retrying cannot succeed and just burns the run's budget. doc-processor deliberately does not auto-retry any 4xx for the same reason.

And writing back

The write direction is the mirror image. When doc-writer is on, CogniRunner puts a bound uploadUrl and uploadAuthHeader in the prompt; the model passes them to whichever create tool it chose, and the finished file is POSTed straight onto the Jira issue as an attachment. The model is told to pass clientHint: "interactive" so what lands in the Jira comment is a clean one-liner and not a metadata dump.

LimitCaps that apply in this path: 50 MB per file on the doc-processor side (READ_DOC_MAX_BYTES), 25 MB on the upload side (WRITE_DOC_MAX_BYTES, deliberately half, because Forge web-trigger payload limits are tighter), and CogniRunner's own 10 MB per file / 20 MB per validation limits on attachment analysis. The tightest of the three wins.

34What rules are actually worth building?

Four that pay for themselves. Each is one validator or post-function with a plain-English prompt — no code.

Attachment actually says what the ticket claims
A validator on the transition into Review. Prompt: "Read the attached specification. Block this transition if it does not contain an explicit rollback plan, and quote the section you looked at." The user gets a specific reason, not a generic refusal.
Generate the deliverable on transition
A post-function on Done. Prompt: "Produce a PDF summarising this issue: the problem, what changed, and the verification steps. Attach it." With doc-writer on, create-pdf runs and the file lands on the issue.
Contract review before approval
A validator using read-doc in focused mode. Prompt: "Find the payment terms and notice period in the attached contract. Block if the notice period is under 30 days." One question, one answer, minimal context burned.
Fact-check what was submitted
A validator that calls fact-check. It needs the web-search MCP configured as well — the two servers work as a pair — and it returns the claims with sources rather than a verdict you have to take on trust.
TipStart with a read-only validator. Turn doc-writer off until you have watched a few runs in CogniRunner's execution log. A rule that writes files to Jira issues is much harder to walk back than one that only reads.

35What is the hosted demo, and what is it not?

One instance of this exact repo, running on our Mac Studio behind a Tailscale Funnel. It exists so you can evaluate the tools in 60 seconds without installing anything — not so you can build on it.

The endpoint on this page is the server described in the "Run it as a server" part, with the configuration shown there and the secrets removed. There is nothing different about the code — the demo is a deployment, not a product tier.

Hosted demoSelf-hosted
Time to first tool callAbout a minuteAbout five minutes
Where created files landOur disk, retrieved over a signed linkYour disk, wherever DOC_OUTPUT_DIR points
Rate limitOurs, shared, modestYours
UptimeA machine in an office. No SLA.Yours
Suitable for confidential documentsNoYes
Suitable for productionNoYes
CarefulDo not send anything confidential to the hosted demo. Documents you read are parsed on our machine and documents you create are written to it. It is a demo. If the content matters, the setup guide above takes five minutes and the software is identical.

36How do I connect to the hosted demo?

Request a key with your email, then add one entry to your client config. Three active keys per address; revoke or self-host past that.

  1. 1Use the key form on this page. The site calls the server's provisioning endpoint, which mints a tenant bearer exactly as the admin API does.
  2. 2The bearer is shown once and emailed to you. It is stored on our side only as an argon2 hash — we cannot read it back to you.
  3. 3Add it to your MCP client with the endpoint below.
  4. 4Confirm with a read-doc on a small local PDF before building anything on it.

Claude Code

claude mcp add --transport http doc-processor \
  https://worksmacstudio.tailfc4700.ts.net/docproc/mcp \
  --header "Authorization: Bearer YOUR_DEMO_KEY"

Any client that takes a remote MCP entry

{
  "mcpServers": {
    "doc-processor": {
      "url": "https://worksmacstudio.tailfc4700.ts.net/docproc/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_DEMO_KEY",
        "X-ZAI-Key": "YOUR_OWN_VISION_KEY (optional — only for scanned-PDF OCR)",
        "X-Output-Dir": "my-folder (optional — a folder ON OUR SERVER)"
      }
    }
  }
}
NoteThe hosted server is keyless for OCR by design. No Z_AI_API_KEY is set on it. If you want to OCR a scanned PDF through the demo, pass your own key in X-ZAI-Key — it is used for that request only and is never stored. Every tool except fact-check then needs no key at all; fact-check additionally needs a web-search bearer and a Serper key, which the demo does not supply.
Careful`X-Output-Dir` organises files on our machine, not yours. A remote server writes to the remote disk — that is a client/server boundary, not a setting. Files come back as signed, expiring /files/download links. If you want documents in your own Finder, run it locally over stdio.

37Troubleshooting

The failures that actually happen, with the one command that distinguishes each from its look-alike.

SymptomMost likely causeDo this
No tools in the client at allThe server never launchedRun the stdio handshake probe. If it prints 17 tools, the problem is client config, not the server.
Works in Claude Code, not in Claude Desktop or LM Studionode is not on the GUI app's PATHReplace "command": "node" with the output of which node.
Unexpected token … in JSONSomething wrote to stdout that was not protocolnpm run lint:no-console-log. If you edited src/, this is why.
A scanned PDF returns empty textNo vision key, so OCR is offSet Z_AI_API_KEY (or X-ZAI-Key on the hosted server). Check the value does not contain "api" or "key" — placeholders are rejected on purpose.
Vision API key not configured although it is setThe value contains the substring "api" or "key"Use the real token, or lm-studio for a local endpoint.
OCR 404s against a local endpointZ_AI_BASE_URL is missing its trailing slashhttp://localhost:1234/v1/ — the URL is built by string concatenation.
create-pdf fails, everything else worksChromium was never downloadednpx puppeteer browsers install chrome.
Files appear somewhere unexpectedOutput root is process.cwd() by defaultSet DOC_OUTPUT_DIR. Remember a remote server writes to its own disk.
Duplicate-title refusalThe registry already has that titleSwitch to edit-doc, or pass preventDuplicates: false.
HTTP: every request rejected through the funnel, fine on localhostHost mismatch in DNS-rebinding protectionAdd both the bare host and host:port to ALLOWED_HOSTS, then restart.
HTTP: 404 on /v1/admin/* with a correct tokenWorking as designed — admin is localhost-onlyRun the curl on the server itself.
Tenant tokens vanish after a restartDATA_DIR points inside the repo, or is not writableMove it outside the repo; on systemd add it to ReadWritePaths.
Long OCR calls time out in the clientThe client's per-server timeout is 60 sSet "timeout": 300000 on the server entry if your client supports it.

The three commands worth knowing

# 1. Does the server work at all, with no client involved?
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node src/index.js 2>/dev/null | tail -1

# 2. Is the HTTP service alive?
curl -s localhost:10000/healthz

# 3. Does the bearer work end to end, from outside?
curl -s -X POST https://your-host.tailXXXX.ts.net/docproc/mcp \
  -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
TipRead the logs before changing anything. stdio: the server's stderr, plus logs/server.log in the directory your client launched it from (not the repo). HTTP under launchd: ~/Library/Logs/mcp-doc-processor.log. Under systemd: journalctl --user -u mcp-doc-processor -f. Claude Desktop also keeps its own per-server log at ~/Library/Logs/Claude/mcp-server-doc-processor.log.
NoteStill stuck? The repo is MIT-licensed and the issue tracker is open. A stack trace from the log and the exact mcp.json entry (with the bearer removed) is usually enough to diagnose it in one round trip.

17 tools

Four jobs: read something, make something, change something, or keep track of what you made. The decision table is in the manual.

read-doc

Read a PDF, DOCX, Excel or PowerPoint — summary, in-depth, or one focused question. Local path, or a remote HTTPS URL.

detect-format

Decide the right format and tone before creating anything. Returns a ready-to-use creation plan.

create-doc

A styled, editable Word DOCX — headings, tables, headers, footers, page numbers, 8 presets.

create-pdf

A final, print-and-sign PDF from markdown. Rendered through headless Chromium.

create-markdown

A README, spec or runbook — with an optional generated table of contents and YAML frontmatter.

create-excel

An XLSX workbook: multiple sheets, styled headers, and live formulas.

create-pptx

An editable deck — one slide per '## ' heading, with bullets, tables, charts and speaker notes.

edit-doc

Append or replace in a DOCX. Append patches the XML in place, so formatting and images survive.

edit-excel

Append rows, add a sheet, or replace one sheet's data — keeping the workbook's styles.

edit-pptx

Preview, append slides, or replace a slide in an existing deck.

fact-check

Verify a document's claims against the live web. Calls the web-search MCP per claim and cites sources.

list-documents

Search the registry by category, tag and title — to find a document or catch a duplicate.

list-templates

List the blueprints a create call can be validated against.

dna

Document DNA — project-wide header, footer and style defaults, memories, and usage-driven evolution.

blueprint

Learn, list and delete structural blueprints extracted from documents you liked.

drift-monitor

Fingerprint a document and get told when its structure changes underneath you.

get-lineage

Trace provenance — which sources informed a document, and what was derived from it.

How it works

About 18,000 lines of plain ESM JavaScript, no build step, no framework. The file-by-file tour is in the manual.

Two entry points, one registry

`src/index.js` speaks stdio, `src/server.js` speaks HTTP, and both dispatch through `tool-registry.js`. The transports cannot drift on what the tools are or what they accept.

Parsers in, generators out

`pdf-parse`, `mammoth` and `xlsx` read; `docx`, `pptxgenjs`, `xlsx-js-style` and headless Chromium write. Nothing in the read path touches the network unless you enable OCR.

Document DNA over three levels

System defaults, then project `.document-dna.json`, then user `.document-user.json`, with anything explicit in the call winning. That is why documents come out looking like yours without anyone remembering to ask.

A sandboxed output root

`getOutputRoot()` resolves the per-request header first (sandboxed), then `DOC_OUTPUT_DIR`, then the working directory. A hosted tenant can organise files but cannot escape its base.

Bridges, not integrations

One JSON envelope in each direction lets any authenticated HTTPS endpoint feed files in and take files out. CogniRunner is the reference implementation, not a special case.

A registry that remembers

Every created document is indexed with category, tags and lineage — which is what makes duplicate detection, provenance and drift monitoring possible at all.

Using it from Jira, through CogniRunner

A workflow rule that can open the PDF somebody attached, decide whether it says what it needs to say, and attach a generated report back to the issue.

Two connection paths

The hosted bridge works with every AI provider — CogniRunner dials the MCP's HTTPS URL from Forge. The local path works only with LM Studio, but keeps the documents on your own machine.

Compare them →

Single-use attachment capabilities

CogniRunner mints a one-shot URL and bearer per attachment. The model calls read-doc with url + authHeader; the capability dies on first use, and nothing is copied to a shared disk.

How the flow works →

Your provider never sees the token

Every MCP call is brokered by CogniRunner inside Atlassian Forge. The AI provider receives the tool's result — never the URL, never the bearer, never the raw attachment.

See CogniRunner →
Self-hosting for CogniRunner? Your funnel must answer on port 443. Forge egress reaches a *.ts.net Tailscale Funnel URL on port 443 and nothing else — 8443 and 10000 are blocked, and an arbitrary domain is refused. Publish on a path prefix instead: tailscale funnel --bg --set-path=/docproc 10000. Full walkthrough →
Free demo key

Try it without installing anything

One instance of this exact repo, on our Mac Studio behind a Tailscale Funnel. Endpoint: https://worksmacstudio.tailfc4700.ts.net/docproc/mcp

We email you a link to get the key. The key is for evaluation on a shared, rate-limited demo server that may be reset at any time. Getting it also adds the address to the list that gets an email when a new tutorial goes out; every one of those has a one-click unsubscribe.

The demo is an evaluation surface, not infrastructure. Documents you read are parsed on our machine, documents you create are written to it, the rate limit is shared and there is no SLA. For anything confidential or production-facing, self-host — it takes five minutes and the code is identical.

Related documentation

This server is one half of a pair, and both are wired into the same Jira app.

MCP Web Search

The sibling server. fact-check calls it per claim, so if you want document verification you need both — same setup shape, same launchd and Tailscale recipe.

CogniRunner

The Jira workflow app that consumes this MCP — AI validators, conditions and post-functions, with the MCP setup documented on its own page.

Local AI

Everything we have written about running models on your own hardware — including the LM Studio setup this server's local OCR path uses.

AI advisory

If you would rather have this wired into your own stack than do it yourself — that is the work we do.

Frequently Asked Questions

Do I need an API key to use it?

Mostly not. Fifteen of the seventeen tools need no key at all — PDF, DOCX, XLSX and PPTX are parsed locally with no network access. Two do: read-doc needs a vision key only to OCR a scanned, image-only PDF, which you can point at a local model in LM Studio instead of a cloud provider; and fact-check always needs a bearer for the web-search MCP plus your own Serper key.

Where do the files it creates end up?

In the working directory the client launched the server from, which for a coding agent is your open project. Set DOC_OUTPUT_DIR to pin a fixed location. A remote server writes to the remote host's disk — that is a client/server boundary, not a setting, so run it locally over stdio if you want documents in your own Finder.

Is it really open source?

Yes, MIT licensed, full source on GitHub. Clone it, run it, fork it, ship it in a product. The hosted demo is one deployment of that same repo, not a different build.

What is the difference between the hosted demo and self-hosting?

The code is identical. The demo saves you five minutes of setup and costs you control: files are written to our machine, the rate limit is shared, and there is no SLA. Anything confidential or production-facing should be self-hosted.

Can it edit a Word document without destroying the formatting?

Yes — edit-doc in append mode patches document.xml directly, so the original fonts, headers, footers and embedded images are untouched. Replace mode rebuilds the body and does not preserve formatting, so choose deliberately.

How does it read a Jira attachment?

Through the remote-read bridge. CogniRunner mints a single-use URL and bearer for the attachment and hands them to the model; read-doc fetches the JSON envelope, decodes it into a per-call temp directory, parses it, and deletes the temp directory afterwards. Nothing is copied to a shared disk.

Does it work with clients other than Claude?

Yes. It is a standard MCP server over stdio, so LM Studio, Cursor, Cline, Roo Code, Continue, Qwen Code and any MCP host work with the same command-plus-args config.

What happens to a scanned PDF with no vision key?

It returns empty text rather than an error, because there is no text layer to extract. Check the result rather than assuming a read succeeded — and set a vision key if you need OCR.

MIT licensed — yours to run

Open source and free

Clone it, run it, fork it, ship it. The hosted demo is one deployment of the same repository — there is no paid tier holding anything back.

View on GitHubSet it upJoin the Community