Notes from the work

One email when a tutorial or migration write-up goes live. Nothing else, and one click to leave.

LeanZero

Two people in Romania doing Atlassian migrations, Forge apps and practical AI work for teams that would rather talk to the person doing the job. Most of what we learn ends up on this site.

Services

  • Atlassian Migrations
  • Atlassian FastShift
  • Atlassian Maintenance
  • Forge App Development
  • AI Development Consultation

Topics

  • Jira
  • Jira Service Management
  • Confluence
  • Bitbucket
  • Atlassian Forge
  • Cloud Migration
  • Local AI
  • AI Coding
  • Certifications
  • Atlassian Team
  • All topics

Company

  • Blog
  • Tutorials
  • Contact

Community

  • Join Discord
  • Support this site

© 2026 LeanZero. All rights reserved.

Privacy PolicyTerms of Service24/7 SupportTrust CenterSecurityDPALegal notice
  1. Home
  2. Portfolio
  3. Mcp Web Search
MCP Server, Open Source, MIT

MCP Web Search

Set it up from scratch in five minutes, then give any AI client live search, page, PDF, GitHub and OpenAPI reading.

Set it upView on GitHub
36-section manual 11 tools Node · TypeScript

Set it up in five minutes

Node 20 or newer, git, ~260 MB of disk, and a Serper key (2,500 free credits, no card). Unlike its sibling this one is TypeScript, so the build step is not optional.

1

Clone, install, build

terminal
git clone https://github.com/leanzero-srl/mcp-web-search.git
cd mcp-web-search
npm install
npm run build      # REQUIRED — creates dist/index.js
pwd                # copy this — the client config needs the absolute path

Cannot find module .../dist/index.js is what you get when you skip the build. It is the most common first-run failure, and the error names the file rather than the cause.

2

Get a key, point your client at it

Claude Code
claude mcp add web-search --scope user \
  -e SERPER_API_KEY=your_key_here \
  -e USE_SERPER_ONLY=true \
  -- node /ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js
everything else — mcp.json
{
  "mcpServers": {
    "web-search": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js"],
      "timeout": 120000,
      "env": {
        "SERPER_API_KEY": "your_key_here",
        "USE_SERPER_ONLY": "true"
      }
    }
  }
}
3

Prove it works — then prove the key works

terminal
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node dist/index.js 2>/dev/null | tail -1

# A healthy server answers with all 11 tools.
Eleven tools is not proof the search works. The server registers its tools whether or not the key is valid, so a missing SERPER_API_KEY looks perfectly healthy until the first real search comes back empty. Run one live search before you trust the setup.

Prerequisites, the four search engines compared honestly, whether you need the Playwright browsers, every client, and what each failure means are all in part 1 of the manual.

Which way should I run it?

Three deployments, one codebase. Most people want the first, and the manual documents all three end to end.

Local, over stdio

Recommended for almost everyone

Your MCP client launches the server as a child process. No port, no listener, and your search queries go to the engine and nowhere else.

  • Clone, install, build
  • Nothing on the network
  • Your Serper key stays local
Install it

An always-on HTTPS service

When something remote has to reach it

The HTTP transport under launchd or systemd, published through a Tailscale Funnel with per-tenant bearer tokens. The recipe our own Mac Studio runs.

  • launchd / systemd files
  • Tailscale Funnel, no open ports
  • Keyless: callers bring their own key
Run the server

The hosted demo

To evaluate it in a minute

One instance of this exact repo on our Mac Studio. Free key by email — and you still bring your own Serper key, so every search is billed to you.

  • Free key by email
  • No install at all
  • Not for production
Get a demo key

Claude Code

One `claude mcp add` command. Non-agent cohort — `claude-code` is not on the whitelist — so disk-writing tools embed content inline.

Claude Desktop

One entry in `claude_desktop_config.json` — with `"timeout": 120000`, which is not optional.

LM Studio

Non-agent cohort, so content is embedded inline. Also the bridge Forge apps reach through.

Cursor · Cline · Roo · Qwen

Same block everywhere. Qwen Code's 30-second default is the most-reported issue and the timeout is the fix.

Atlassian Forge apps

CogniRunner exposes four of the eleven tools, brokered from inside Forge.

The manual

Every question this server raises, answered in order: install it, connect your client, run it as a service the way we do, tune it, use all 11 tools, read the code, wire it into Jira through CogniRunner, and fix it when it misbehaves. 36 sections.

Written against the server's own source and against the live deployment behind the hosted demo — the launchd, Tailscale and hardening steps are the configuration actually running, with the secrets removed.

In this manual
  • What do I need before I start?
  • How do I install it?
  • How do I get a search key?
  • Do I need the Playwright browsers?
  • How do I prove the install actually works?
  • How do I add it to Claude Code?
  • How do I add it to Claude Desktop?
  • How do I add it to LM Studio?
  • How do I add it to Cursor, Cline, Roo Code, or Qwen Code?
  • The server does not show up, or every search comes back empty. What now?
  • Do I want stdio or HTTP?
  • How do I start the HTTP transport?
  • How do I keep it running on macOS? (launchd)
  • How do I keep it running on Linux? (systemd)
  • How do I publish it over HTTPS without opening a port?
  • How do I issue and revoke access tokens?
  • What must I check before this faces the internet?
  • How do I make it fast?
  • How do I make the results better?
  • Why do I get a file path instead of content?
  • Every environment variable, in one place
  • Which tool do I call?
  • How do I run a good search?
  • How do I extract from something specific?
  • What happened to the file it said it saved?
  • What is in the repo, file by file?
  • What actually happens on one search?
  • How do I add my own tool or search engine?
  • How do I know I have not broken it?
  • What does CogniRunner do with this MCP?
  • How do I connect it over the hosted bridge?
  • How do I connect it locally through LM Studio?
  • What rules are actually worth building?
  • What is the hosted demo, and what is it not?
  • How do I connect to the hosted demo?
  • Troubleshooting

01What do I need before I start?

Node 20 or newer, git, ~260 MB of disk, and a search key. There is a build step here — this one is TypeScript, unlike its sibling, and skipping it is the most common first-run failure.

MCP Web Search is a Node.js server that gives an AI agent live access to the web: search, full-page extraction, PDF reading, GitHub repository crawling, OpenAPI spec discovery, and multi-round progressive research. It speaks the Model Context Protocol over stdio, so your client launches it as a child process — and it also has an HTTP transport for the remote case.

Requirements

WhatVersionWhyCheck it
Node.js20 or newer (we run 24)package.json declares engines: { node: ">=20" }.node --version
npm10 or newerInstalls and builds.npm --version
gitanyCloning the repo.git --version
Disk~260 MB, or ~1 GB with browsersMeasured on a clean clone: 256 MB of node_modules. Playwright's browsers are separate and optional — Chromium plus its headless shell is ~520 MB, and adding Firefox (the other default in BROWSER_TYPES) takes it past 750 MB.df -h .
A Serper key2,500 free creditsThe default search path, and in the stock configuration the only one. Sign-up takes ~2 minutes and needs no card, but the free grant is one-time, not monthly. Going without it means enabling browser fallbacks — see below.serper.dev
NoteThis is a fork. The upstream project is mrkrsl/web-search-mcp, MIT licensed, and this fork keeps that licence. What was added: the orchestration layer, the enterprise guardrails, the HTTP transport with tenant auth, client-aware output shaping, the semantic cache, and the specialised GitHub / OpenAPI / PDF extractors.

02How do I install it?

Clone, install, build. The build is not optional — the entry point your client launches is dist/index.js, and it does not exist until you run it.

The whole install

git clone https://github.com/leanzero-srl/mcp-web-search.git
cd mcp-web-search
npm install
npm run build      # tsc + esbuild -> dist/index.js and dist/bundle.cjs
pwd                # copy this — the client config needs the absolute path

What the build produces

  • `dist/index.js` — the stdio entry point. This is what goes in args in your client config.
  • `dist/http-server.js` — the HTTP transport, for the always-on server case.
  • `dist/bundle.cjs` — a single-file CommonJS bundle, useful when you want to drop the server somewhere without a node_modules tree.
Careful`Cannot find module .../dist/index.js` means you skipped `npm run build`. It is the single most common first-run failure, and the error names the file rather than the cause. npm run dev (tsx watch) runs from source if you are iterating on the code.

Playwright is a dependency, but `npm install` does NOT download the browsers — verified on a clean clone: the published playwright package registers no install script, and launching with an empty browser cache fails with browserType.launch: Executable doesn't exist at …. If you intend to use browser fallbacks you must run npx playwright install chromium firefox as a separate step. You do not need them for the default Serper-only configuration — see "Do I need the Playwright browsers?" below.

03How do I get a search key?

Sign up at serper.dev for 2,500 free credits and paste the key into one environment variable. The grant is one-time, not monthly — and there are keyless engines too, which are worse. Here is the honest comparison.

  1. 1Go to serper.dev and sign up. No card is required, and you get 2,500 credits — enough to evaluate the server thoroughly. Be clear-eyed that this is a one-time grant, not a monthly allowance: once it is spent, searching costs money (credit packs start around $50 for 50,000 credits, valid six months) or you fall back to the keyless engines.
  2. 2Copy the API key from the dashboard.
  3. 3Put it in SERPER_API_KEY in the env block of your MCP client config.
  4. 4Leave USE_SERPER_ONLY at its default of true unless you specifically want browser fallbacks.

The four engines

EngineNeeds a keyAlso needsHonest assessment
SerperYes — 2,500 free, then paidNothingThe right choice, and the path the defaults take. A real search API: no scraping, no bot detection, no CAPTCHA, and fastest by a wide margin.
DuckDuckGoNoUSE_SERPER_ONLY=false, a raised per-tool timeout, and browsers for its fallback legTried second. Uniquely, it starts with a plain Axios fetch of html.duckduckgo.com and only escalates to a browser — so it is the one keyless engine that can technically answer with no browsers at all, though the page it returns is often unparseable.
BingNoUSE_SERPER_ONLY=false, a raised per-tool timeout, and the Playwright browsers (~750 MB)Tried third, through a headless browser. Works until it does not — expect intermittent blocking.
BraveNoUSE_SERPER_ONLY=false, a raised per-tool timeout, and the Playwright browsers (~750 MB)Tried last, same machinery.
CarefulThere is no `SEARCH_ENGINE` setting — the order is hard-coded. getEnginePriorityOrder() in src/browser-engine.ts returns ['api','webkit','chromium','firefox'], which maps to Serper → DuckDuckGo → Bing → Brave. The only thing you configure is whether the non-Serper engines run at all (USE_SERPER_ONLY / ENABLE_BROWSER_FALLBACKS). Any guide telling you to set SEARCH_ENGINE=brave — including earlier versions of this page — is describing a variable the code never reads.
Careful"Keyless" does not mean "no setup". Bing and Brave are driven through a headless Playwright browser, so reaching them needs USE_SERPER_ONLY=false AND the browser binaries. Worse, the per-tool timeouts cut them off: measured on this machine with browsers installed and no key, the DuckDuckGo leg alone burned 10.4 s against an 8,000 ms TOOL_TIMEOUT_SEARCH_SUMMARIES, logging Search timeout reached before attempting engine chromium — Bing and Brave were never tried. Raising it to 120,000 ms produced results, at a quality score of 0.15, below the default RELEVANCE_THRESHOLD of 0.3. Keyless is a research setting, not a deployment.
CarefulA missing search key does not announce itself. With the default USE_SERPER_ONLY=true and no working key, search-engine.ts returns { results: [], engine: "serper-failed-no-fallback" } and the tool answers "No results found for … Try different or broader keywords, or use full-web-search for a deeper crawl." That is indistinguishable from a genuinely empty search. Verified against the live hosted server. If every query comes back "no results", suspect the key first.
TipFailover between engines is real, but only among the engines you have actually enabled. With browser fallbacks on, a blocked or rate-limited engine hands off to the next in priority order. Note that PARALLEL_SEARCH (default true) is a misnomer — the source renamed the function to searchWithPriorityFallbacks precisely because it is sequential, trying engines in priority order and stopping at the first that clears the quality bar.
NoteA `GITHUB_TOKEN` is optional but worth it. Without one, get-github-repo-content uses the unauthenticated GitHub API and its 60-requests-per-hour limit, which one repo crawl can exhaust. A read-only personal access token raises that to 5,000.

04Do I need the Playwright browsers?

Not if you have a Serper key — the default configuration never launches one, and skipping them saves about 750 MB. They become close to mandatory the moment you try to run keyless.

Extraction runs in two stages. Stage 1 is a fast HTTP fetch with Axios and works for the large majority of pages. Stage 2 spins up a headless browser to get past bot detection and to render JavaScript-only pages. Stage 2 is off by default (USE_SERPER_ONLY=true) because it is dramatically slower and heavier, and because for most research the Stage-1 result is the same result.

CarefulThis is a separate step — `npm install` will not do it for you. Verified on a clean clone. Skip it and every browser fallback dies with browserType.launch: Executable doesn't exist at …, which names a path rather than telling you to install anything.

Turn browser fallbacks on

# Install the browsers (once) — REQUIRED, npm install does not do this:
npx playwright install chromium firefox

# Then, in your client's env block:
#   "USE_SERPER_ONLY": "false"
#   "BROWSER_HEADLESS": "true"
#   "BROWSER_TYPES": "chromium,firefox"
#   "MAX_BROWSERS": "3"

Turn them on when

  • You have no Serper key. Bing and Brave are browser-driven, so without the browsers, USE_SERPER_ONLY=false and a raised TOOL_TIMEOUT_*, there is effectively no search path. (DuckDuckGo alone tries a plain HTTP fetch first, but rarely returns anything parseable.)
  • A site you need consistently returns near-empty content — a sign the real content is rendered client-side.
  • You are extracting from something behind Cloudflare or a similar interstitial.
  • You are running unattended research where a slow answer beats no answer.
CarefulBrowsers cost memory and ~750 MB of disk, and the memory cost compounds. MAX_BROWSERS (default 3) times CONTEXT_POOL_SIZE (default 10) is your worst case. On a small VPS, leave fallbacks off — the OOM killer arriving mid-research is a much worse failure than an empty extraction.

05How do I prove the install actually works?

Drive the protocol from a shell and count the tools. Eleven means it is genuinely working, before any client is involved.

Handshake and list the tools

printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node dist/index.js 2>/dev/null \
  | tail -1 | python3 -c 'import json,sys; print(len(json.load(sys.stdin)["result"]["tools"]), "tools")'
TipExpected: 11 tools. If you get that, the install is sound and every remaining problem is client configuration or a search key.

Then prove the search path end to end

npm test                    # the vitest suite
npm run test:unit           # units only, fastest
npm run test:integration    # hits the real network — needs SERPER_API_KEY
CarefulA tool list of 11 with a missing or invalid `SERPER_API_KEY` still looks healthy. The server registers its tools regardless; the failure surfaces only on the first search, as empty results or an engine error. Run one real search before you trust the setup.

06How do I add it to Claude Code?

One command, with the Serper key passed as an environment variable. Point at dist/index.js, not src/.

Add it

claude mcp add web-search --scope user \
  -e SERPER_API_KEY=your_key_here \
  -e USE_SERPER_ONLY=true \
  -- node /ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js
  • Everything after -- is the command Claude Code runs. The -- is required.
  • --scope user puts it in every project. --scope project writes a committed .mcp.json; --scope local (the default) is this project only.
  • Verify with claude mcp list, then run one real search. A registered-but-keyless server looks identical to a working one until you search.
  • claude mcp remove web-search undoes it.
TipClaude Code is NOT on the agentic whitelist. It reports clientInfo.name = "claude-code", which prefix-matches no entry in AGENTIC_NAMES (claude-ai does not match it), so it falls to the default and gets the non-agent shape: disk-writing tools embed the content inline, capped by MAX_OUTPUT_LENGTH. If you would rather have the compact path-only shape, add "claude-code" to AGENTIC_NAMES in src/client-detect.ts and rebuild. See "Why do I get a file path instead of content?".

07How do I add it to Claude Desktop?

Edit one JSON file, fully quit the app, reopen. Use an absolute node path — the GUI does not inherit your shell PATH.

Where the config lives

OSPath
macOS~/Library/Application Support/Claude/claude_desktop_config.json
Windows%APPDATA%\\Claude\\claude_desktop_config.json
Linux~/.config/Claude/claude_desktop_config.json

claude_desktop_config.json

{
  "mcpServers": {
    "web-search": {
      "command": "/usr/local/bin/node",
      "args": ["/ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js"],
      "timeout": 120000,
      "env": {
        "SERPER_API_KEY": "your_key_here",
        "USE_SERPER_ONLY": "true",
        "MAX_CONTENT_LENGTH": "100000"
      }
    }
  }
}
Careful`"timeout": 120000` is not decoration. A full search with content extraction takes 30–90 seconds. Clients that default to 30 or 60 seconds kill the call mid-flight and report a timeout that looks like a server fault. Milliseconds, not seconds.
  1. 1Merge your entry into the existing mcpServers object — do not replace the file, or you silently delete every other server.
  2. 2Use the absolute path from which node for command.
  3. 3Fully quit the app (Cmd-Q / tray → Quit) and reopen. Closing the window does not reload the config.
  4. 4Ask for a live search on something that changed this week. A model answering from training data looks the same as a working search until you check the dates.

08How do I add it to LM Studio?

Same mcp.json shape — and this is the path CogniRunner and other Forge apps depend on, so it is worth getting exactly right.

LM Studio mcp.json

{
  "mcpServers": {
    "web-search": {
      "command": "/usr/local/bin/node",
      "args": ["/ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js"],
      "timeout": 120000,
      "env": {
        "SERPER_API_KEY": "your_key_here",
        "USE_SERPER_ONLY": "true"
      }
    }
  }
}
  1. 1Open Program → Edit mcp.json in LM Studio.
  2. 2Add the entry above with the absolute path to dist/index.js.
  3. 3Save — LM Studio reloads MCP servers on save, no restart needed.
  4. 4Load a model that supports tool calling. Without tool support the server is connected and never called.
NoteLM Studio is a non-agent client, and the server knows it. It has no sibling filesystem MCP, so tools that would normally return a file path embed the content inline instead — capped by MAX_OUTPUT_LENGTH. That detection is a prefix match on clientInfo.name in src/client-detect.ts, and it defaults to non-agent, so an unknown client gets the safe inline shape rather than an unreadable path.
CarefulIf CogniRunner will use this, the entry must be named `web-search` exactly — which is what the block above already uses, so nothing to change. CogniRunner looks the integration up by that key, and any other name means the tools are never offered to the model, with no error to tell you why. Its sibling is the one that trips people: the doc-processor server has to be registered as doc-reader.

09How do I add it to Cursor, Cline, Roo Code, or Qwen Code?

Identical block, different file. The only thing to remember everywhere is the timeout.

ClientWhere the config livesCohort
Claude Codeclaude mcp addNon-agent (inline content) — claude-code is not on the whitelist
Cursor.cursor/mcp.json, or Settings → MCPNon-agent (inline content)
Cline (VS Code)cline_mcp_settings.jsonAgentic (file paths)
Roo CodeThe extension's MCP panelAgentic
Continue~/.continue/config.jsonAgentic
Qwen Code.qwen/settings.jsonNon-agent
Anything elseAny mcpServers object with command + args + envNon-agent by default

The universal block

{
  "mcpServers": {
    "web-search": {
      "command": "node",
      "args": ["/ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js"],
      "timeout": 120000,
      "trust": true,
      "env": {
        "SERPER_API_KEY": "your_key_here",
        "USE_SERPER_ONLY": "true",
        "MAX_CONTENT_LENGTH": "100000",
        "EXTRACT_CONCURRENCY": "3"
      }
    }
  }
}
TipQwen Code's default is about 30 seconds and produces Request timed out (-32001) on any real search. It is the most frequently reported issue against this server and "timeout": 120000 is the entire fix.

10The server does not show up, or every search comes back empty. What now?

Two different faults with two different tests. Distinguish them first — treating a key problem as a connection problem wastes an afternoon.

If no tools appear at all

  1. 1Run the stdio probe from the setup part. 11 tools means the server is fine and the problem is the client.
  2. 2Did you build? ls dist/index.js. No build, no entry point.
  3. 3Is `node` findable by the client? Use the absolute path from which node. GUI apps do not inherit your shell PATH.
  4. 4Is the JSON valid? python3 -m json.tool < your-config.json. A trailing comma disables the whole file in most clients.
  5. 5Did you fully restart? Claude Desktop and Cursor need a real quit; LM Studio reloads on save.

If tools appear but searches return nothing

  1. 1Recognise the symptom. With the default USE_SERPER_ONLY=true, a missing or invalid key produces "No results found for … Try different or broader keywords" — not an authentication error. It looks exactly like a genuinely empty search, so check the key before you rewrite the query.
  2. 2Check the key reached the process. A key in your shell is not a key in the client's environment — it has to be in the config's env block.
  3. 3Test the key directly: curl -s -X POST https://google.serper.dev/search -H "X-API-KEY: $SERPER_API_KEY" -H "Content-Type: application/json" -d '{"q":"test"}' | head -c 200.
  4. 4Look for the circuit breaker. After SERPER_BREAKER_FAILURES (default 5) consecutive failures the breaker opens for SERPER_BREAKER_COOLDOWN_MS (default 30 s) and every search fails fast until it closes.
  5. 5Try the keyless engines as a control: USE_SERPER_ONLY=false, the Playwright browsers installed (npx playwright install chromium firefox), and a raised TOOL_TIMEOUT_SEARCH_SUMMARIES — at its 8,000 ms default the browser engines are cut off before they are reached. There is no SEARCH_ENGINE variable to set. If results then appear, the problem is definitely the Serper key.
  6. 6Check the relevance filter. RELEVANCE_THRESHOLD (default 0.3) discards low-scoring content. On a niche query it can filter everything; set ENABLE_RELEVANCE_CHECKING=false to confirm that is what is happening.
CarefulEmpty results are not always a fault. The quality scorer is doing its job when it drops a page of navigation chrome. Prove it with ENABLE_RELEVANCE_CHECKING=false before changing anything else — if results appear, tune RELEVANCE_THRESHOLD rather than turning the filter off for good.

11Do I want stdio or HTTP?

stdio if the agent runs on your machine. HTTP if a Forge app, a phone or claude.ai has to reach it. The HTTP transport is the fork's addition — upstream has none.

stdio (dist/index.js)HTTP (dist/http-server.js)
Who launches itYour MCP client, as a child processlaunchd / systemd, as a service
Network exposureNoneA listener, plus whatever fronts it
AuthNone neededPer-tenant argon2 bearers; optional OAuth 2.1 resource server
Where the search key livesYour client's env blockThe server's env, or per-request via X-Serper-Key
Who can use itOne machineAnything that reaches the URL
NoteUpstream deliberately has no HTTP transport, on the reasoning that HTTP means a listener, an auth surface, TLS, rate limiting and DDoS exposure — and that the LM Studio bridge avoids all of it. That reasoning is sound. This fork added HTTP anyway because a Forge app on a non-LM-Studio provider has no other route, and it carries the auth, rate limiting and host checking that the argument demands.

What follows is the configuration actually running on the Mac Studio behind this site's hosted demo, secrets removed.

12How do I start the HTTP transport?

node dist/http-server.js with a handful of variables, bound to loopback. /healthz first, always.

Start it by hand first

cd /path/to/mcp-web-search
PORT=8443 \
DATA_DIR="$HOME/Library/Application Support/mcp-web-search" \
USE_SERPER_ONLY=true \
TENANT_RATE_LIMIT=20 \
MCP_WEB_SEARCH_ADMIN_TOKEN="$(openssl rand -base64 32)" \
node dist/http-server.js

curl -s localhost:8443/healthz     # -> {"ok":true,...}

Server environment variables

VariableExampleWhat it does
PORT8443Listen port.
PUBLIC_HOSThost.tailXXXX.ts.net:8443DNS-rebinding protection — the Host header must match. Include the port when the proxy forwards a non-default one.
ALLOWED_HOSTS(rarely needed)Extra accepted Host values, comma-separated. Not needed for the bare-host / host:port pair — http-server.ts derives the bare host from PUBLIC_HOST itself. Use it only for a genuinely different hostname, such as a custom domain. (Its sibling doc-processor does not do this, so there the second form really is required.)
DATA_DIR~/Library/Application Support/mcp-web-searchtenants.json and the insights log — not the cached documents. Keep it outside the repo so git clean cannot delete your tenants.
OUTPUT_DIR / CRAWL_CACHE_DIRunder the working directoryWhere saved research markdown and OpenAPI specs actually land. Unset, they default to docs/research-output under the cwd and docs/technical beside the repo — both inside a `git clean`'s reach. Set them if saved documents matter.
MCP_WEB_SEARCH_ADMIN_TOKENopenssl rand -base64 32Guards /v1/admin/*.
TENANT_RATE_LIMIT20Requests per minute per tenant. Lower than doc-processor's because a search costs far more.
FILE_TOKEN_SECRET(random)Signs /files/download links. Set it explicitly so rotating other secrets does not invalidate live links.
OAUTH_ISSUER / OAUTH_JWKS_URL / OAUTH_AUDIENCE(blank)Turns the server into an OAuth 2.1 resource server for claude.ai web. Blank = bearer-only, which is the posture we run.

Endpoints

RouteMethodPurpose
/healthzGETLiveness. Unauthenticated.
/mcpPOSTThe MCP endpoint. Bearer required.
/files/downloadGETSigned, expiring link to a research markdown or OpenAPI spec the tools saved.
/v1/admin/tenantsGET / POST / DELETEMint, list, rotate, revoke. Localhost-only — a proxied request gets 403.
/v1/provisionPOSTSelf-service minting, reachable through the funnel but guarded by a shared PROVISION_SECRET. This is what the key form on this page calls.
TipRun the server keyless. Set no SERPER_API_KEY on it and let each caller bring their own via the X-Serper-Key header (or ?serper_key=). That is how our instance runs: a compromise of the host leaks no search credential, and every tenant's usage is billed to their own key. X-GitHub-Token and X-Output-Dir work the same way.

13How do I keep it running on macOS? (launchd)

The same three files as its sibling: a plist with no secrets, a wrapper that sources a mode-0600 env file, and exec node.

run-http-server.sh

#!/bin/bash
# Sources server.env so secrets stay out of the launchd plist, then execs node.
# launchd's KeepAlive restarts this on crash; `exec` keeps node as the job's PID.
set -euo pipefail

APP_DIR="$HOME/Library/Application Support/mcp-web-search"
REPO="/Users/you/Projects/mcp-web-search"

set -a
# shellcheck disable=SC1091
source "$APP_DIR/server.env"
set +a

exec /usr/local/bin/node "$REPO/dist/http-server.js"

server.env

# NO global SERPER_API_KEY — each tenant brings their own per request
# (X-Serper-Key header or ?serper_key query).
PORT=8443
PUBLIC_HOST=your-host.tailXXXX.ts.net:8443
DATA_DIR="/Users/you/Library/Application Support/mcp-web-search"
USE_SERPER_ONLY=true
TENANT_RATE_LIMIT=20
MCP_WEB_SEARCH_ADMIN_TOKEN=<openssl rand -base64 32>
FILE_TOKEN_SECRET=<openssl rand -base64 32>

~/Library/LaunchAgents/com.user.mcp-web-search.plist

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
  "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>
    <string>com.user.mcp-web-search</string>

    <key>ProgramArguments</key>
    <array>
        <string>/bin/bash</string>
        <string>/Users/you/Library/Application Support/mcp-web-search/run-http-server.sh</string>
    </array>

    <key>WorkingDirectory</key>
    <string>/Users/you/Projects/mcp-web-search</string>

    <key>RunAtLoad</key>
    <true/>
    <key>KeepAlive</key>
    <true/>

    <key>StandardOutPath</key>
    <string>/Users/you/Library/Logs/mcp-web-search.log</string>
    <key>StandardErrorPath</key>
    <string>/Users/you/Library/Logs/mcp-web-search.log</string>
</dict>
</plist>

Install, start, watch

# The directory does not exist yet — create it before writing the two files into it.
mkdir -p ~/Library/Application\ Support/mcp-web-search

chmod 600 ~/Library/Application\ Support/mcp-web-search/server.env
chmod +x  ~/Library/Application\ Support/mcp-web-search/run-http-server.sh

launchctl load -w ~/Library/LaunchAgents/com.user.mcp-web-search.plist
launchctl list | grep mcp-web-search
curl -s localhost:8443/healthz
tail -f ~/Library/Logs/mcp-web-search.log
CarefulDo not load the job before `server.env` exists and is filled in. set -euo pipefail makes the wrapper exit immediately on a missing source, KeepAlive restarts it, and you get a tight crash loop that fills the log file. Unload, fix, reload.

14How do I keep it running on Linux? (systemd)

A user unit with an EnvironmentFile, plus the browser caveat that catches people who enabled Playwright.

~/.config/systemd/user/mcp-web-search.service

[Unit]
Description=MCP Web Search (HTTP transport)
Wants=network-online.target
After=network-online.target

[Service]
Type=simple
WorkingDirectory=%h/Projects/mcp-web-search
EnvironmentFile=%h/.config/mcp-web-search/server.env
ExecStart=/usr/bin/node %h/Projects/mcp-web-search/dist/http-server.js
Restart=always
RestartSec=3

NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ReadWritePaths=%h/.local/share/mcp-web-search

[Install]
WantedBy=default.target

Install and verify

mkdir -p ~/.config/mcp-web-search ~/.local/share/mcp-web-search

# Write server.env into ~/.config/mcp-web-search/ — and on Linux it MUST set a
# Linux DATA_DIR. The macOS example above uses an "Application Support" path; left
# unset, DATA_DIR falls back to the working directory, so tenants.json is written
# into the repo where a git clean can delete it (verified in a container).
#   DATA_DIR=/home/you/.local/share/mcp-web-search
#   OUTPUT_DIR=/home/you/.local/share/mcp-web-search
# then:
chmod 600 ~/.config/mcp-web-search/server.env
loginctl enable-linger $USER        # survive logout
systemctl --user daemon-reload
systemctl --user enable --now mcp-web-search
journalctl --user -u mcp-web-search -f
curl -s localhost:8443/healthz
CarefulEvery path in `ReadWritePaths=` must already exist, or systemd refuses to start the unit — which is why the install block creates them before daemon-reload.
Careful`PrivateTmp=true` plus Playwright is a trap. Headless browsers want a real writable temp area and a browser cache. If you have enabled browser fallbacks, either drop PrivateTmp, or set PLAYWRIGHT_BROWSERS_PATH to a directory listed in ReadWritePaths. Symptom: searches work until one falls back to a browser, then that one call fails.

15How do I publish it over HTTPS without opening a port?

Tailscale Funnel, on a path prefix, sharing one hostname with the doc-processor. Real certificate, no port forwarding, no dynamic DNS.

  1. 1Install and join. brew install tailscale (or your distro package), sudo tailscale up, then tailscale status for the MagicDNS name.
  2. 2Enable MagicDNS and HTTPS certificates in the admin console under DNS.
  3. 3Grant the funnel attribute in Access Controls: "nodeAttrs": [{ "target": ["autogroup:member"], "attr": ["funnel"] }].
  4. 4Publish it. tailscale funnel --bg --set-path=/websearch 8443 → https://your-host.tailXXXX.ts.net/websearch/mcp.
  5. 5Match `PUBLIC_HOST` / `ALLOWED_HOSTS` to what clients actually send, then restart.
  6. 6Verify from off-network — a phone on cellular, not the LAN.

One funnel, both MCP servers

tailscale funnel --bg --set-path=/websearch 8443     # this server
tailscale funnel --bg --set-path=/docproc   10000    # doc-processor
tailscale funnel --bg --set-path=/lmstudio  1234     # LM Studio, for the Forge bridge
tailscale funnel status
CarefulPath routing on `:443` sends a bare `Host`; a dedicated funnel port sends `host:port`. This server accepts both automatically from PUBLIC_HOST, so if funnel requests are rejected the cause is almost always an unset or mistyped PUBLIC_HOST — not a missing ALLOWED_HOSTS. Symptom: /healthz fine on localhost, every funnel request rejected.
NoteTailnet-only instead? tailscale serve has the same syntax and keeps the service reachable only from your own devices. Bearer auth still applies. Choose it unless a Forge app or claude.ai web genuinely needs to reach in.

16How do I issue and revoke access tokens?

One bearer per consumer, argon2-hashed, minted over a localhost-only admin API — with a rotation grace window so a rotation does not break a running client.

Mint, list, rotate, revoke

ADMIN=<MCP_WEB_SEARCH_ADMIN_TOKEN>

curl -s -X POST localhost:8443/v1/admin/tenants \
  -H "Authorization: Bearer $ADMIN" -H "Content-Type: application/json" \
  -d '{"displayName":"alice-laptop"}'

curl -s            localhost:8443/v1/admin/tenants              -H "Authorization: Bearer $ADMIN"
curl -s -X POST    localhost:8443/v1/admin/tenants/<id>/rotate  -H "Authorization: Bearer $ADMIN"
curl -s -X DELETE  localhost:8443/v1/admin/tenants/<id>         -H "Authorization: Bearer $ADMIN"
  • The admin surface is localhost-only. Anything arriving with an X-Forwarded-For — i.e. through the funnel — is refused with 403, even with a valid admin token.
  • Rotation keeps a grace window during which the old bearer still works, so you can update a client without a hard cutover.
  • Tokens are argon2 hashes in tenants.json. A stolen file yields no working bearers, and a lost bearer cannot be recovered — rotate instead.
  • `/v1/provision` is the self-service door: reachable through the funnel, but guarded by a shared PROVISION_SECRET that only your own backend knows. It is what the demo-key form on this page calls.

17What must I check before this faces the internet?

A search server is a more attractive target than a document server — it makes outbound requests on a stranger's behalf. Nine checks.

  1. 1Bind to loopback. lsof -nP -iTCP -sTCP:LISTEN | grep 8443 should show 127.0.0.1, not *.
  2. 2`PUBLIC_HOST` set and verified through the funnel, not from localhost.
  3. 3A strong `MCP_WEB_SEARCH_ADMIN_TOKEN`. openssl rand -base64 32.
  4. 4Confirm the admin surface is closed from outside: the same curl through the funnel must return 403.
  5. 5A strong `PROVISION_SECRET`, or none at all if you are not running self-service minting.
  6. 6`chmod 600` on `server.env` and `tenants.json`; DATA_DIR outside the repo.
  7. 7`TENANT_RATE_LIMIT` set low. 20/min is our number. A search costs far more than a document read, and this is the limit that stops one tenant becoming your whole Serper bill.
  8. 8Run keyless. No SERPER_API_KEY on the server; callers bring X-Serper-Key. Otherwise every tenant spends your quota.
  9. 9Log rotation, or KeepAlive plus an unrotated log eventually fills the disk.
CarefulThis server fetches URLs that callers choose. That is server-side request forgery by design, and the SSRF guard is what keeps it honest — it refuses localhost, 127.0.0.1 and RFC1918 targets. Do not weaken it, and do not run this on a host that sits inside a network where reaching an internal address would matter.
LimitWhat it does not do: cap total egress, cap total Serper spend, or persist an audit trail of what each tenant searched beyond the logs. TENANT_RATE_LIMIT bounds the rate, not the month. Watch the Serper dashboard.

18How do I make it fast?

Five settings account for nearly all the latency. The defaults are already the fast configuration — this is what to change when they are not enough.

The settings that actually move the needle

SettingFastThoroughWhat you are trading
USE_SERPER_ONLYtrue (default)falseBrowser fallbacks — and with them the entire keyless-engine path, since Bing/Brave/DuckDuckGo are browser-driven. true means seconds and Serper only; false means tens of seconds, plus client-side-rendered pages and the keyless engines.
PARALLEL_SEARCHfalsetrue (default)Despite the name, not concurrent. true picks searchWithPriorityFallbacks, which tries engines in priority order and returns the first that clears the quality threshold; false picks the plain sequential path. Only reachable at all when browser fallbacks are on, so with the default Serper-only configuration it does nothing.
EXTRACT_CONCURRENCY53 (default)How many pages are fetched at once. Higher is faster until you trip a rate limit.
MAX_CONTENT_LENGTH50000100000 (default)Characters per page. Lower is faster and cheaper in context; higher keeps more of a long document.
SEMANTIC_CACHE_ENABLEDtrue (default)falseReusing near-identical earlier searches. Off means always-fresh results and always-full latency.
Note`MAX_CONTENT_LENGTH` was lowered from 500 KB to 100 KB for a measured reason: ten results at 500 KB each is 5 MB of strings per search, and the garbage-collection pressure from that dominated the wall-clock time. If you raise it, watch memory as well as latency.

Per-tool wall-clock budgets (milliseconds)

VariableDefault
TOOL_TIMEOUT_FULL_SEARCH18000
TOOL_TIMEOUT_SEARCH_SUMMARIES8000
TOOL_TIMEOUT_SINGLE_PAGE12000
TOOL_TIMEOUT_PDF12000
TOOL_TIMEOUT_GITHUB18000
TOOL_TIMEOUT_OPENAPI15000
TOOL_TIMEOUT_PROGRESSIVE20000
CarefulEvery default sits under ~20 seconds on purpose. The Forge → LM Studio → MCP chain has to complete inside Forge's ~25-second function timeout. Raise these and a CogniRunner rule starts failing at the Forge layer with an error that says nothing about this server. If you do not use the Forge path, raise them freely.

19How do I make the results better?

Quality is a filter, not a search. The scorer decides what reaches your model — and it can be too strict as easily as too loose.

VariableDefaultWhat it does
ENABLE_RELEVANCE_CHECKINGtrueTurns the quality scorer on. Off means everything extracted reaches the model, chrome included.
RELEVANCE_THRESHOLD0.3Minimum score (0.0–1.0) for content to survive. Raise for precision, lower for recall.
MIN_CONTENT_LENGTH200Bytes below which a result is treated as empty — kills cookie-banner-only extractions.
MAX_OUTPUT_LENGTH50000Inline-content budget for the two disk-writing tools on non-agent clients. The response-wide cap is MAX_TOOL_RESPONSE_CHARS (250,000).
SEMANTIC_CACHE_TTL3600000One hour. Lower it when freshness matters more than latency.

How to diagnose a bad result

  • Empty results on a niche query — likely the threshold. Set ENABLE_RELEVANCE_CHECKING=false once as a control; if content appears, lower RELEVANCE_THRESHOLD to about 0.15 rather than leaving the filter off.
  • Results that are just navigation and cookie banners — the opposite problem. Raise the threshold, and check MIN_CONTENT_LENGTH is not too low.
  • Stale results on a fast-moving topic — the semantic cache is answering. A response with engine: semantic-cache tells you so outright. Lower the TTL or disable the cache for that workload.
  • Right pages, missing content — the site is client-rendered. This is the case that justifies USE_SERPER_ONLY=false.
TipThe semantic cache is opt-in infrastructure, not just a flag. It needs UPSTASH_VECTOR_REST_URL and UPSTASH_VECTOR_REST_TOKEN; without them initialisation is skipped and the server logs that it skipped. If you expected caching and never see a semantic-cache hit, that is why.

20Why do I get a file path instead of content?

Because your client was identified as agentic. Two tools change shape based on who is calling, and the rule is a prefix match you can read in one file.

CohortClientsWhat the disk-writing tools return
AgenticCline, Claude Desktop, Roo Code, Continue, claude-aiJust the file path. These clients run a sibling filesystem MCP and read the file themselves — embedding the content would waste the context twice.
Non-agentLM Studio, MCP Inspector, anything unrecognisedThe full content inline, truncated to MAX_OUTPUT_LENGTH. No filesystem access, so a path would be a dead end.

The two tools affected are research_and_save_to_markdown and get-openapi-spec. The rule lives in src/client-detect.ts as a case-insensitive prefix match against clientInfo.name from the initialize handshake, and it defaults to non-agent — an unknown caller gets the safe inline shape rather than an unreachable path.

The behaviour is verified end to end: calling get-openapi-spec against the Petstore spec as Cline returns 788 bytes with no inline spec; the same call from a non-agent name returns the spec embedded.

TipYour client is agentic but has no filesystem tool? Add its name to AGENTIC_NAMES in src/client-detect.ts and rebuild — or remove it from the list to force the inline shape. The comment above that array asks you to whitelist conservatively, because a false positive silently hands LM Studio users file paths they cannot open.

21Every environment variable, in one place

The full reference. Nothing here is required — every one has a working default except the search key.

Search and API

VariableDefaultPurpose
SERPER_API_KEY(none)The Serper-first fast path. Effectively required for good results.
SEARCH_ENGINE—Not a real setting. No code reads it; the engine order is hard-coded in getEnginePriorityOrder(). Listed here only because older guides mention it.
USE_SERPER_ONLYtrueSkip Playwright fallbacks. The performance default.
ENABLE_BROWSER_FALLBACKSfalseLegacy alias; true overrides USE_SERPER_ONLY.
PARALLEL_SEARCHtrueMisnomer — selects searchWithPriorityFallbacks, which is sequential priority-order failover. false selects the older sequential path. Only reachable when browser fallbacks are on.
SEARCH_ENGINE_MAX_RPM50Requests per minute to the engine.
SEARCH_ENGINE_RESET_MS60000RPM window.
SERPER_BREAKER_FAILURES5Consecutive failures before the circuit breaker opens.
SERPER_BREAKER_COOLDOWN_MS30000How long it stays open before a probe.
DEBUG_BING_SEARCHfalseVerbose Bing parsing logs.

Extraction and quality

VariableDefaultPurpose
MAX_CONTENT_LENGTH100000Characters kept per page.
MIN_CONTENT_LENGTH200Below this, treat as empty.
DEFAULT_TIMEOUT6000Per-extraction timeout, ms.
EXTRACT_CONCURRENCY3Parallel page extractions per search.
BROWSER_FALLBACK_THRESHOLD3Failures before escalating to a browser.
ENABLE_RELEVANCE_CHECKINGtrueQuality scoring on/off.
RELEVANCE_THRESHOLD0.3Minimum score to keep content.

Browser and stealth

VariableDefaultPurpose
BROWSER_HEADLESStruefalse shows the browser — useful once, when debugging an extraction.
BROWSER_TYPESchromium,firefoxWhich browsers the pool may launch.
MAX_BROWSERS3Concurrent browser instances.
CONTEXT_POOL_SIZE10Browser contexts kept warm.
CONTEXT_MAX_AGE60000How long a context lives, ms.
CONTEXT_REUSE_TIMEOUT20000Wait before reusing a context.
USE_LEGACY_POOLfalseThe older, more conservative pool implementation.

Guardrails, GitHub and cache

VariableDefaultPurpose
MAX_REQUESTS_PER_MINUTE30Currently inert — parsed at start-up, but its limiter's checkAndRecord() has no call sites.
MAX_REQUESTS_PER_SECOND10Currently inert — its globalThrottler singleton is never imported. The only enforced limit is TENANT_RATE_LIMIT on the HTTP transport.
MAX_OUTPUT_LENGTH50000Inline-content budget for the two disk-writing tools on non-agent clients.
MAX_TOOL_RESPONSE_CHARS250000The response-wide cap every tool answer passes through.
GITHUB_TOKEN(none)Raises the GitHub API limit from 60/h to 5,000/h.
GITHUB_MAX_DEPTH3Directory depth when crawling a repo.
GITHUB_MAX_FILES50Files extracted per repo.
SEMANTIC_CACHE_ENABLEDtrueCache near-identical searches.
SEMANTIC_CACHE_MAX_SIZE1000Cached results retained.
SEMANTIC_CACHE_TTL3600000Cache lifetime, ms.
UPSTASH_VECTOR_REST_URL / _TOKEN(none)Required for the semantic cache to initialise at all.

22Which tool do I call?

Cheapest that answers the question, always. The gap between get-web-search-summaries and full-web-search is roughly an order of magnitude in both time and context.

ToolCostCall it when
get-web-search-summariesCheapestA quick fact check. Snippets and descriptions, no page extraction.
full-web-searchStandardReal research. Top results with full, cleaned Markdown content. The semantic cache sits in this path.
progressive-web-searchMost expensiveA hard or ambiguous question. Query expansion across multiple rounds and engines.
get-single-web-page-contentCheapYou already have the URL. Targeted extraction, no search.
get-pdf-contentCheapThe URL ends in .pdf. Text extraction with a browser fallback.
get-github-repo-contentVariesReading a repository. crawl walks it, list shows one directory, file returns one file.
get-openapi-specCheapDiscovering and downloading an OpenAPI/Swagger spec, JSON or YAML.
get-website-sitemapCheapMapping a site. keywords filters; extractTopMatching also extracts the top N in the same call.
research_and_save_to_markdownExpensiveMulti-URL research written to a markdown file per result.
list-cached-documentsFreeWhat did I save earlier? Filter all / openapi / research.
read-cached-documentFreeRead one of those back by name. Path-traversal characters are refused.
TipThe decision rule that matters most: if a snippet would answer it, do not run a full search. Most agents default to full-web-search for everything and spend ten times the latency and context for the same answer.
Note`get-website-sitemap` collapses a three-call chain. Sitemap → filter → extract used to be three round-trips; keywords plus extractTopMatching (1–5) does it in one. filter-sitemap-urls was folded into it and no longer exists as a separate tool.

23How do I run a good search?

Pick the right tool, then let the orchestration do its job. Understanding the six stages tells you which knob to reach for when a result disappoints.

What happens on one full-web-search

  1. 1Intent detection. query-intent-detector.ts classifies the query — does this need a summary or real content?
  2. 2Engine selection. Serper first when a key is present. With USE_SERPER_ONLY=true (the default) that is the only attempt — a failure returns no results rather than falling through. With browser fallbacks on, Bing, Brave and DuckDuckGo are tried in priority order, stopping at the first that clears the quality bar.
  3. 3Stage 1 extraction, fast. Axios fetches each result; Cheerio and the enhanced extractor turn the HTML into clean Markdown.
  4. 4Stage 2 extraction, stealth — only if enabled and Stage 1 failed. A pooled Playwright browser renders the page and gets past bot detection.
  5. 5Quality scoring. Each extraction is scored for relevance and noise; anything under RELEVANCE_THRESHOLD is dropped before it reaches your model.
  6. 6Caching. The result is written to the semantic cache, so a near-identical query later returns instantly with engine: semantic-cache.

The three search shapes

// Cheapest — snippets only
{ "query": "Atlassian Forge function timeout limit" }

// Standard — full extracted content
{ "query": "Atlassian Forge function timeout limit",
  "limit": 5,
  "includeContent": true }

// Hardest questions — query expansion, several rounds
{ "query": "why do Forge async events sometimes run twice" }
Careful`progressive-web-search` is genuinely expensive. It expands the query into synonyms, related concepts and intent-based variants, then searches each. That is the right tool for a question a single query cannot express — and the wrong tool for "what is the capital of France".

24How do I extract from something specific?

Four specialised extractors, each with a mode or parameter that decides how much you pay. Choosing wrong is what makes a repo crawl take a minute.

get-single-web-page-content
One URL, cleaned Markdown out. Use it whenever the prompt already names the page — searching for a URL you have is pure waste.
get-pdf-content
Only for `.pdf` URLs, not HTML pages that mention a PDF. Extracts readable text, with a browser fallback for PDFs served behind interstitials.
get-github-repo-content
Three modes. crawl (the default) walks the repo and returns the README plus per-file previews — previewLength defaults to 500 characters and caps at 5,000. list returns one directory. file with path returns one file in full. Reach for `file` or `list` first; crawl on a large repo is the expensive call, bounded by GITHUB_MAX_DEPTH and GITHUB_MAX_FILES.
get-openapi-spec
Discovers and downloads an OpenAPI/Swagger spec in JSON or YAML. For non-agent clients the spec is also embedded inline up to OPENAPI_INLINE_CAP (50 KB) so a model with no filesystem can still read it.

GitHub, cheap to expensive

// Cheapest: one file you already know you need
{ "url": "https://github.com/org/repo", "mode": "file", "path": "src/index.ts" }

// Cheap: what is in this directory?
{ "url": "https://github.com/org/repo", "mode": "list", "path": "src" }

// Expensive: walk the whole thing
{ "url": "https://github.com/org/repo", "mode": "crawl", "previewLength": 500 }
CarefulWithout a `GITHUB_TOKEN` you get 60 API requests per hour, unauthenticated. One crawl of a medium repo can spend most of that, and the next call fails for an hour with a message about rate limiting that says nothing about the crawl that caused it.

25What happened to the file it said it saved?

research_and_save_to_markdown and get-openapi-spec write to disk. Where that disk is depends entirely on whether the server is local or remote — and the two cached-document tools are how you get the content back either way.

  1. 1`list-cached-documents` with category: "all" | "openapi" | "research" shows what is there.
  2. 2`read-cached-document` returns one by file name, inline. It refuses any name containing path-traversal characters.
  3. 3On the HTTP transport, the tools also return a signed /files/download link so a remote caller can fetch the file over HTTPS.
CarefulA remote server saves to the remote disk. Running against the hosted demo, a "saved" research file is on our Mac Studio, not in your Downloads folder. Read it back with read-cached-document or the signed link. If you want files on your own machine, run the server locally over stdio.
NoteWhether you get content or a path is decided by your client's cohort, not by these tools — see "Why do I get a file path instead of content?". LM Studio gets the content inline; Cline gets the path and reads it with its filesystem MCP.

26What is in the repo, file by file?

About 11,000 lines of TypeScript across 30 files in src/. Two entry points, one tool registration, and a clean split between the search layer and the extraction layer.

The tree

mcp-web-search/
├─ src/
│  ├─ index.ts                     # ENTRY: stdio transport
│  ├─ http-server.ts               # ENTRY: Express HTTP transport
│  ├─ server.ts                    # The 11 tools: schemas + handlers. Start here.
│  ├─ auth.ts                      # argon2 tenant bearers, /v1/admin, /v1/provision
│  ├─ oauth.ts                     # OAuth 2.1 resource-server validation
│  ├─ search-engine.ts             # Serper / Bing / Brave / DuckDuckGo + failover
│  ├─ progressive-search-engine.ts # Multi-round expansion
│  ├─ semantic-expander.ts         # Synonyms, related concepts, intent variants
│  ├─ query-intent-detector.ts     # Summary or deep extraction?
│  ├─ browser-engine.ts            # Playwright Stage 2
│  ├─ browser-pool.ts              # Browser lifecycle
│  ├─ context-pool.ts              # Browser-context reuse
│  ├─ enhanced-content-extractor.ts# HTML -> clean Markdown
│  ├─ content-quality-scorer.ts    # Relevance and noise scoring
│  ├─ github-extractor.ts          # crawl / list / file
│  ├─ openapi-extractor.ts         # Spec discovery, JSON + YAML
│  ├─ pdf-extractor.ts             # PDF text, with browser fallback
│  ├─ semantic-cache.ts            # Upstash-backed near-duplicate cache
│  ├─ crawl-cache.ts               # Short-lived page cache
│  ├─ request-deduplicator.ts      # Collapses identical in-flight requests
│  ├─ enterprise-guardrails.ts     # Per-tool argument allow-lists, output caps
│  ├─ rate-limiter.ts              # Global + per-tenant limits
│  ├─ client-detect.ts             # Agentic vs non-agent cohort
│  ├─ request-context.ts           # Per-request key / token / output dir
│  ├─ download-registry.ts         # Signed /files/download links
│  ├─ observability.ts, insights.ts, logger.ts, utils.ts, types.ts
└─ tests/  unit/ integration/ e2e/

The files that carry the weight

FileWhat lives there
src/server.tsRead this first. All 11 tool schemas, descriptions and handlers, registered through registerToolsOn(target) so both transports share one definition.
src/search-engine.tsEngine selection, the Serper circuit breaker, failover, parallel search.
src/enhanced-content-extractor.tsThe two-stage extraction and the HTML-to-Markdown conversion. Where result quality is actually decided.
src/content-quality-scorer.tsThe filter that keeps chrome out of your context — and the one to look at when good results go missing.
src/enterprise-guardrails.tsOutput caps and the session 404 circuit-breaker. It also defines per-tool argument allow-lists, but validateToolArgs() has no callers, so argument shape is enforced only by each tool's Zod schema in server.ts.
src/client-detect.tsThe agentic whitelist. Small, consequential, well commented.
src/request-deduplicator.tsCollapses identical in-flight requests into one. Free latency when an agent asks the same thing twice.
Tip`registerToolsOn(target)` is the design decision worth noticing. Both index.ts (stdio) and http-server.ts call it against their own server instance, so the two transports can never drift on what tools exist or what they accept.

27What actually happens on one search?

Eleven steps from JSON-RPC frame to Markdown. Almost every configuration question answers itself once you can point at the step it belongs to.

  1. 1Transport receives the frame. HTTP additionally runs requireAuth (argon2 bearer against tenants.json) and the per-tenant rate limiter.
  2. 2Request context is established — X-Serper-Key, X-GitHub-Token and X-Output-Dir are lifted off the request into request-context.ts so nothing downstream needs them threaded through.
  3. 3The semantic cache is consulted first — before any rate limiting or concurrency overhead. A near-match returns immediately, tagged engine: semantic-cache.
  4. 4Only on a miss, the deduplicator checks for an identical in-flight request and joins it rather than starting a second.
  5. 5Intent is detected — summary or deep extraction.
  6. 6The engine is chosen and queried, with the Serper circuit breaker in play — and, only if browser fallbacks are enabled, priority-order failover to Bing, Brave and DuckDuckGo.
  7. 7Stage 1 extraction — Axios plus Cheerio, EXTRACT_CONCURRENCY pages at a time.
  8. 8Stage 2 extraction if enabled and needed — a pooled Playwright browser.
  9. 9Quality scoring drops anything under the threshold.
  10. 10The response is shaped by the client cohort, capped at MAX_OUTPUT_LENGTH, and written back to the cache.
NoteThe code calls `console.log` freely — `src/server.ts` alone has 59. What keeps stdout clean is a shim, not discipline: src/index.ts calls installStdioSafeConsoleShim() before anything else, redirecting console.* to stderr. That is what makes stdout a safe protocol channel. http-server.ts deliberately skips the shim, since HTTP does not multiplex over stdout.

28How do I add my own tool or search engine?

A tool is one target.tool(...) call in server.ts plus a guardrail entry. An engine is one branch in search-engine.ts. Neither needs a framework.

Adding a tool

  1. 1Add a `target.tool(...)` block in `src/server.ts` — name, description, a Zod schema for the arguments, and the handler. Both transports pick it up automatically.
  2. 2Optional: add an allow-list entry to `src/enterprise-guardrails.ts` for when validateToolArgs() is wired up. Today only the Zod schema is enforced, so the schema is what actually guards the tool.
  3. 3Decide the client-cohort behaviour. If the tool writes to disk, follow the pattern in research_and_save_to_markdown: path for agentic clients, inline content for everyone else.
  4. 4Respect the timeout budget. Add a TOOL_TIMEOUT_* variable and keep the default under ~20 s if the Forge path matters to you.
  5. 5Add a test under tests/unit/ and, if it touches the network, tests/integration/.
  6. 6`npm run build`, then restart your client.

Adding a search engine

  1. 1Add the branch in src/search-engine.ts alongside Serper, Bing, Brave and DuckDuckGo.
  2. 2Return the shared result shape from src/types.ts — everything downstream depends on it.
  3. 3Wire it into the priority order in getEnginePriorityOrder() — that array is the only place engine selection happens; there is no SEARCH_ENGINE variable.
  4. 4Give it a rate limit. SEARCH_ENGINE_MAX_RPM is global; a scraped engine usually needs its own, stricter one.
CarefulTool descriptions are the model's routing logic. Look at how the existing ones state cost explicitly — "cheapest", "only when summaries are insufficient", "ONLY for .pdf URLs". Without that, a model reaches for full-web-search every time and your latency triples for no gain.

29How do I know I have not broken it?

Vitest, in three tiers. Only the middle tier needs the network, so the fast tier is genuinely fast.

Run them

npm test                  # everything
npm run test:unit         # units only — no network, fastest
npm run test:integration  # hits the real web; needs SERPER_API_KEY
npm run test:coverage     # with coverage
npm run lint              # eslint over src/**/*.ts

What to run when

  • Changed extraction or scoring? test:unit first, then test:integration — the second is the only one that proves a real page still extracts.
  • Changed a tool schema? test:integration — tests/integration/ is the only suite that drives real tool calls. Note tests/e2e/ exists but is empty, so npm run test:e2e exits 1 with 'No test files found'.
  • Changed auth or the HTTP transport? The whole suite. This is the surface facing the internet.
  • Changed `client-detect.ts`? test:unit — classifyClientNameAsAgentic is exported specifically so tests exercise the production rule rather than a copy of it.
Careful`test:integration` failing is not automatically your fault. It queries the live web: an engine can be having a bad afternoon, or your Serper quota can be spent. Re-run once, check the Serper dashboard, and only then suspect the change.

30What does CogniRunner do with this MCP?

It lets a Jira workflow rule check the live web while it runs — verify a claim, read a linked page, fetch a referenced PDF — instead of reasoning from a model's training cutoff.

CogniRunner is LeanZero's AI workflow app for Jira: validators, conditions and post-functions on transitions. With web-search connected, a validator can go and check whether the thing the ticket claims is actually true, and say where it looked.

The four tools CogniRunner exposes, in priority order

ToolWhen the model should reach for it
get-web-search-summariesFirst. Quick fact checks. Cheapest by a wide margin.
get-single-web-page-contentOnly when the prompt names a specific URL.
get-pdf-contentOnly for `.pdf` URLs — not HTML pages that mention a PDF.
full-web-searchOnly when summaries are genuinely insufficient.
CarefulWeb search is slow in this path — 30 to 90 seconds. CogniRunner's own guidance budgets at most one call per validation unless the prompt explicitly demands more. A rule that fans out to five searches will hit the Forge function timeout, and the user sees a failed transition with no useful explanation.
NoteThe other seven tools are deliberately not exposed. A curated four-tool surface routes better than eleven, and the repo-crawling and file-writing tools have no meaningful role inside a workflow transition.

31How do I connect it over the hosted bridge?

Service URL, Tenant Bearer, and — because this MCP is keyless — a Serper key. The port requirement is the thing that catches self-hosters.

  1. 1Get a URL and bearer. A free demo key from this page, or your own self-hosted instance with a minted tenant token.
  2. 2Get a Serper key from serper.dev — 2,500 free credits on sign-up, no card, one-time. The MCP is keyless by design, so the key is supplied per tenant rather than baked into the server, and each tenant's searches are billed to their own key.
  3. 3Open CogniRunner's admin panel → Settings → MCP Integrations and turn on the web-search card.
  4. 4Paste the Service URL (must be https://). Hosted demo: https://worksmacstudio.tailfc4700.ts.net/websearch/mcp.
  5. 5Paste the Tenant Bearer — minimum 16 characters, masked once saved, preserved when you edit only the URL.
  6. 6Paste the Serper key. Optionally a GitHub token too, if you want repo reading.
  7. 7Run the card's probe before building a rule on it.
CarefulSelf-hosting? Your funnel must answer on port 443. CogniRunner's Forge egress allow-list permits a *.ts.net Tailscale Funnel URL on port 443 (and context7's own host) and nothing else — 8443 and 10000 are blocked, and an arbitrary domain like https://mycompany.com/mcp is refused. The allow-list ships inside the installed app and cannot be changed by an installer.
TipPath prefixes are the way through that. tailscale funnel --bg --set-path=/websearch 8443 publishes a loopback service on https://<host>.ts.net/websearch — port 443 outside, 8443 inside. Add the bare hostname to ALLOWED_HOSTS as well as the host:port form, because the two routes send different Host headers.

32How do I connect it locally through LM Studio?

An mcp.json entry named exactly web-search, with a timeout of 120000. Both are hard requirements, and neither fails loudly.

CarefulThe entry name must be exactly `web-search`, and "timeout": 120000 is required. CogniRunner finds the integration by that key — any other name and the tools are simply never offered. And LM Studio's default timeout kills a 30–90 second search mid-flight.

LM Studio mcp.json

{
  "mcpServers": {
    "web-search": {
      "command": "/usr/local/bin/node",
      "args": ["/ABSOLUTE/PATH/TO/mcp-web-search/dist/index.js"],
      "timeout": 120000,
      "env": {
        "SERPER_API_KEY": "<YOUR_SERPER_KEY>",
        "USE_SERPER_ONLY": "true"
      }
    }
  }
}

In the local path the Serper key lives here, in the LM Studio host's environment — not in CogniRunner. That is the privacy argument for this route: the query leaves your machine to Serper and to nowhere else, and no LeanZero infrastructure is in the path at all.

Then, on the CogniRunner side

  1. 1Set the provider to LM Studio and give CogniRunner your LM Studio's Tailscale Funnel URL. HTTPS, on *.ts.net, never localhost — Forge cannot reach a private address.
  2. 2Turn on "Serve on Local Network" in LM Studio's developer settings so the funnel relay can reach it.
  3. 3Switch on "Run locally via LM Studio (mcp.json)" in the web-search card.
  4. 4Make sure every other enabled MCP is also local — see the warning below.
  5. 5Load a tool-capable model and run the probe.
CarefulDo not mix local and hosted across enabled MCPs. LM Studio cannot combine native local plugins and hosted-bridge function tools in one request. Mix them and CogniRunner routes all of them through the hosted bridge, so the "local" one then also needs a Service URL and Bearer — or it quietly stops working.

33What rules are actually worth building?

Four that earn their latency. Each is one validator or post-function with a plain-English prompt.

Verify a vendor claim before approval
A validator on the transition into Approved. Prompt: "The ticket claims this library is still maintained. Check the web and block the transition if the last release is more than 18 months old. Cite what you found." One get-web-search-summaries call, a specific answer.
Check a referenced URL is real and says what is claimed
A validator using get-single-web-page-content. Prompt: "The description links to a documentation page. Read it and confirm it documents the endpoint named in the summary." No search cost — the URL is already known.
Enrich a security ticket with current advisories
A post-function on creation. Prompt: "Search for current advisories for the component named in this issue and add them as a comment with sources." Runs after the transition, so latency does not block anybody.
Fact-check an attached document
The cross-MCP path: doc-processor's fact-check calls this server per claim. It needs both MCPs configured, and on the hosted bridge it needs the web-search bearer and Serper key present in the same context.
TipPrefer post-functions over validators for anything slow. A validator makes a human wait at a transition screen for up to 90 seconds; a post-function does the same work after the transition has already succeeded. Reserve validators for the cases where blocking is genuinely the point.

34What is the hosted demo, and what is it not?

One instance of this exact repo on our Mac Studio behind a Tailscale Funnel, so you can evaluate the tools in a minute. It is an evaluation surface, not infrastructure.

Hosted demoSelf-hosted
Time to first searchAbout a minuteAbout five minutes
Search keyYours, sent per request as X-Serper-KeyYours, in the server env or per request
Rate limit20 requests/minute per tenant, shared machineYours
UptimeA machine in an office. No SLA.Yours
Where saved research landsOur disk; read it back over the tools or a signed linkYour disk
Suitable for productionNoYes
NoteThe hosted server is keyless by design. No SERPER_API_KEY is set on it, so every search is billed to the key you send. Nothing is stored: the key is read from the request, used for that call, and discarded.
CarefulYour queries are search queries. They reach Serper (or whichever engine answers) and they pass through our machine on the way. For anything sensitive, self-host — the setup above is five minutes and the code is identical.

35How do I connect to the hosted demo?

Request a key with your email, add one entry, and send your own Serper key in a header. Three active keys per address.

  1. 1Use the key form on this page. It calls the server's /v1/provision endpoint, which mints a tenant bearer exactly as the admin API does.
  2. 2The bearer is shown once and emailed. It is stored only as an argon2 hash — we cannot read it back.
  3. 3Get a Serper key at serper.dev if you do not have one — 2,500 free credits on sign-up, no card. The hosted server is keyless, so without this header no search can succeed.
  4. 4Add the entry below, then run one real search before building anything on it.

Claude Code

claude mcp add --transport http web-search \
  https://worksmacstudio.tailfc4700.ts.net/websearch/mcp \
  --header "Authorization: Bearer YOUR_DEMO_KEY" \
  --header "X-Serper-Key: YOUR_SERPER_KEY"

Any client that takes a remote MCP entry

{
  "mcpServers": {
    "web-search": {
      "url": "https://worksmacstudio.tailfc4700.ts.net/websearch/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_DEMO_KEY",
        "X-Serper-Key": "YOUR_SERPER_KEY",
        "X-GitHub-Token": "ghp_... (optional — raises the GitHub API limit)"
      }
    }
  }
}
CarefulNo Serper key means no search on the hosted demo. The server runs USE_SERPER_ONLY=true, so there is no keyless fallback: a request without X-Serper-Key returns "No results found …" rather than an authentication error, which reads exactly like an empty search. Two minutes at serper.dev is the whole fix — 2,500 free credits, no card, no obligation afterwards.

36Troubleshooting

The failures that actually happen, each with the command that separates it from its look-alike.

SymptomMost likely causeDo this
Cannot find module .../dist/index.jsYou did not buildnpm run build.
No tools appear at allThe server never launchedRun the stdio probe. 11 tools means the problem is the client.
Works in Claude Code, not Claude Desktop or LM Studionode is not on the GUI app's PATHUse the absolute path from which node.
Request timed out (-32001)The client's per-server timeout is too short"timeout": 120000. Milliseconds. This is the most reported issue on the repo.
Tools appear, every search is emptyNo usable search keyTest the key with a direct Serper curl. Then check it is in the client's env block, not just your shell.
Searches fail fast, all at onceThe Serper circuit breaker is openFive consecutive failures opens it for 30 seconds. Fix the key, then wait out SERPER_BREAKER_COOLDOWN_MS.
Results are staleThe semantic cache is answeringA response tagged engine: semantic-cache confirms it. Lower SEMANTIC_CACHE_TTL or disable the cache.
Good pages, empty contentThe site renders client-side, and browser fallbacks are offUSE_SERPER_ONLY=false plus npx playwright install chromium firefox — both, because BROWSER_TYPES defaults to chromium,firefox.
Everything filtered out on a niche queryThe relevance thresholdENABLE_RELEVANCE_CHECKING=false once as a control, then lower RELEVANCE_THRESHOLD to about 0.15.
GitHub tools rate-limitedUnauthenticated: 60 requests/hourSet GITHUB_TOKEN, and prefer mode: "file" or "list" over "crawl".
A file path where you expected contentYour client is on the agentic whitelistWorking as designed. Read the file, or edit AGENTIC_NAMES in src/client-detect.ts.
HTTP: rejected through the funnel, fine on localhostPUBLIC_HOST unset or mistypedFix PUBLIC_HOST. The bare-host form is accepted automatically here, so ALLOWED_HOSTS is only for a genuinely different hostname.
HTTP: 403 on /v1/admin/*Working as designed — localhost onlyRun the curl on the server itself.
browserType.launch: Executable doesn't exist at …The Playwright browsers were never installed — npm install does not fetch themnpx playwright install chromium firefox. This is a separate, required step whenever USE_SERPER_ONLY=false.
Browser fallbacks fail under systemdPrivateTmp=true and no writable browser cacheDrop PrivateTmp, or point PLAYWRIGHT_BROWSERS_PATH at a ReadWritePaths directory.

The four commands worth knowing

# 1. Does the server work at all?
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | node dist/index.js 2>/dev/null | tail -1

# 2. Is the search key itself good?
curl -s -X POST https://google.serper.dev/search \
  -H "X-API-KEY: $SERPER_API_KEY" -H "Content-Type: application/json" \
  -d '{"q":"test"}' | head -c 200

# 3. Is the HTTP service alive?
curl -s localhost:8443/healthz

# 4. Does the bearer work from outside?
curl -s -X POST https://your-host.tailXXXX.ts.net/websearch/mcp \
  -H "Authorization: Bearer YOUR_KEY" -H "X-Serper-Key: YOUR_SERPER_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
TipRead the logs first. stdio: the server's stderr. HTTP under launchd: ~/Library/Logs/mcp-web-search.log. Under systemd: journalctl --user -u mcp-web-search -f. DEBUG_BING_SEARCH=true adds verbose parsing output when a scraped engine is the suspect.

11 tools, priced by what they cost you

The gap between a summary and a full search is roughly an order of magnitude in both latency and context. Most agents reach for the expensive one every time. The decision table is in the manual.

get-web-search-summaries

Cheapest

Snippets and descriptions, no page extraction. The right tool for a quick fact check.

full-web-search

Standard

Top results with full, cleaned Markdown content. The semantic cache sits in this path.

progressive-web-search

Expensive

Query expansion across multiple rounds and engines, for questions one query cannot express.

get-single-web-page-content

Cheap

Targeted extraction from a URL you already have. No search cost at all.

get-pdf-content

Cheap

Text from a PDF URL, with a browser fallback for PDFs behind interstitials.

get-github-repo-content

Varies

Three modes — crawl the repo, list one directory, or return a single file in full.

get-openapi-spec

Cheap

Discovers and downloads OpenAPI / Swagger specs, JSON and YAML.

get-website-sitemap

Cheap

Maps a site, filters by keyword, and can extract the top N matches in the same call.

research_and_save_to_markdown

Expensive

Multi-URL research written to a markdown file per result.

list-cached-documents

Free

What did I save earlier? Filter by all, openapi or research.

read-cached-document

Free

Reads a saved document back inline. Refuses path-traversal characters.

How one search actually runs

Six stages between your query and the Markdown. Almost every tuning question answers itself once you can point at the stage it belongs to. The full eleven-step trace is in the manual.

Intent detection

The query is classified before anything is searched — does this need a summary, or real extracted content? That decision alone is most of the latency difference between tools.

Engine selection with failover

Serper first when a key is present. On the default Serper-only configuration that is the only attempt — a failure returns no results rather than falling through. With browser fallbacks enabled, a blocked or rate-limited engine hands off in the hard-coded priority order (Serper → DuckDuckGo → Bing → Brave), and a circuit breaker stops hammering one that is down.

Stage 1 — fast HTTP

Axios and Cheerio fetch and clean each result into Markdown, `EXTRACT_CONCURRENCY` pages at a time. This handles the large majority of the web.

Stage 2 — stealth browser

Only if enabled and Stage 1 failed: a pooled Playwright browser renders the page and gets past bot detection. Off by default because it is dramatically slower.

Quality scoring

Every extraction is scored for relevance and noise, and anything under the threshold is dropped before it reaches your model. This is also why a niche query can come back empty.

Semantic caching

A near-identical query later returns instantly, tagged `engine: semantic-cache`. Needs Upstash credentials to initialise — without them it silently skips.

Using it from Jira, through CogniRunner

A workflow rule that checks the live web while it runs — instead of reasoning from a model's training cutoff.

Four tools, in cost order

Summaries first, then a named URL, then a PDF, and full-web-search only when summaries are genuinely insufficient. The other seven tools are deliberately not exposed.

What gets exposed →

Two connection paths

The hosted bridge works with every AI provider. The local path works only with LM Studio — but then the query leaves your machine to the search engine and to nowhere else.

Connect it →

Budget one call per rule

Search takes 30 to 90 seconds in this path. Every per-tool timeout default sits under ~20 s so the Forge → LM Studio → MCP chain fits inside Forge's ~25-second function limit.

See CogniRunner →
Self-hosting for CogniRunner? Your funnel must answer on port 443, and the mcp.json entry must be named exactly web-search. Forge egress reaches a *.ts.net URL on port 443 and nothing else; a differently-named entry means the tools are never offered, with no error to tell you why. Full walkthrough →
Free demo key

Try it without installing anything

One instance of this exact repo, on our Mac Studio behind a Tailscale Funnel. Endpoint: https://worksmacstudio.tailfc4700.ts.net/websearch/mcp

We email you a link to get the key. The key is for evaluation on a shared, rate-limited demo server that may be reset at any time. Getting it also adds the address to the list that gets an email when a new tutorial goes out; every one of those has a one-click unsubscribe.

The hosted server is keyless by design. No search key is set on it, so every search is billed to the Serper key you send in X-Serper-Key. It is read from the request, used for that call, and discarded.
Your queries pass through our machine. For anything sensitive, self-host — five minutes, and the code is identical.

Related documentation

This server is one half of a pair, and both are wired into the same Jira app.

MCP Doc Processor

The sibling server. Its fact-check tool calls this one per claim — so document verification needs both. Same setup shape, same launchd and Tailscale recipe.

CogniRunner

The Jira workflow app that consumes this MCP — AI validators, conditions and post-functions, with the MCP setup documented on its own page.

Local AI

Everything we have written about running models on your own hardware, including the LM Studio bridge this server reaches Forge apps through.

Upstream project

mrkrsl/web-search-mcp — the MIT-licensed project this is forked from. Credit where it is due.

Frequently Asked Questions

Do I need a paid API key?

Not to get started. Signing up at serper.dev takes about two minutes, needs no card, and gives you 2,500 credits — by far the best result quality available. That grant is one-time rather than monthly, so sustained use eventually costs money. Going fully keyless is possible but not free of effort: Bing, Brave and DuckDuckGo are driven through a headless browser, so they need USE_SERPER_ONLY=false and the ~500 MB of Playwright browsers installed. On the stock configuration they are not reachable at all.

Why does every search time out?

Your client's per-server timeout is too short. A full search with content extraction takes 30 to 90 seconds; most clients default to 30 or 60. Set "timeout": 120000 — milliseconds, not seconds. It is the most frequently reported issue against this server.

Why do I get a file path instead of the content?

Because your client was identified as agentic — Cline, Claude Desktop, Roo Code, Continue and claude-ai run a sibling filesystem MCP and read the file themselves, so embedding the content would waste the context twice. Unknown clients default to the safe inline shape.

Why are my results empty when the search clearly should find something?

Two causes, and they look identical. First and most common: no working search key. With the default USE_SERPER_ONLY=true there is no keyless fallback, so the tool answers "No results found — try different or broader keywords" rather than an authentication error. Second: the relevance filter. RELEVANCE_THRESHOLD defaults to 0.3 and can filter everything on a niche query — set ENABLE_RELEVANCE_CHECKING=false once as a control, and if content appears, lower the threshold to about 0.15 rather than leaving the filter off.

Do I need the Playwright browsers?

Not if you have a Serper key — the default configuration never launches one and skipping them saves about 750 MB. They become close to mandatory if you want to run keyless, because Bing and Brave are browser-driven, and you must also raise the per-tool timeouts or the browser engines are cut off before they are reached. Turn them on with USE_SERPER_ONLY=false, also when a site you need consistently returns near-empty content, which is a sign it renders client-side.

Is this the same as the upstream project?

It is a fork of mrkrsl/web-search-mcp, MIT licensed, with the orchestration layer, enterprise guardrails, HTTP transport with tenant auth, client-aware output shaping, semantic cache and the specialised GitHub / OpenAPI / PDF extractors added.

Where do queries actually go?

To the search engine you configured, and to the pages it returns. Self-hosted over stdio, nothing else is in the path. Through the hosted demo, requests also pass through our Mac Studio — so self-host anything sensitive.

Can I use it with CogniRunner in Jira?

Yes, over the hosted bridge with any AI provider, or locally through LM Studio. CogniRunner exposes four of the eleven tools in a deliberate cost order, and the manual documents both paths including the port-443 requirement that catches self-hosters.

MIT licensed — yours to run

Open source and free

Clone it, run it, fork it, ship it. The hosted demo is one deployment of the same repository — there is no paid tier holding anything back.

View on GitHubSet it upJoin the Community