Back to Browse

GetAudio MCP Server

Developer ToolsLow Risk10.0MCP RegistryLocal
Free

Server data from the Official MCP Registry

让 agent 驱动本机的 Verbatim 音视频转写与博主观点分析:提交转写、轮询、搜索转写库。

About

让 agent 驱动本机的 Verbatim 音视频转写与博主观点分析:提交转写、轮询、搜索转写库。

Security Report

10.0
Low Risk10.0Low Risk

Valid MCP server (1 strong, 2 medium validity signals). No known CVEs in dependencies. Package registry verified. Imported from the Official MCP Registry.

8 files analyzed · 1 issue found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

HTTP Network Access

Connects to external APIs or services over the internet.

env_vars

Check that this permission is expected for this type of plugin.

Shell Command Execution

Runs commands on your machine. Be cautious — only use if you trust this plugin.

What You'll Need

Set these up before or after installing:

Verbatim 服务地址,默认 http://127.0.0.1:5001Optional

Environment variable: VERBATIM_URL

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-xyzxinlu-max-verbatim-transcribe-mcp": {
      "env": {
        "VERBATIM_URL": "your-verbatim-url-here"
      },
      "args": [
        "verbatim-transcribe-mcp"
      ],
      "command": "uvx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

Verbatim

Audio → transcript → insight. A local, self-hosted web app that turns hours of audio/video into timestamped text — and then into structured, cross-referenced opinion research.

Built as a personal tool for studying Chinese commentators/KOLs: drop files or paste a channel URL, and Verbatim transcribes every episode, writes a per-episode analysis, and synthesizes one document. The interface is English; the transcribed content keeps its original language.

The app lives in getAudio/ and runs entirely on your machine (Flask, localhost:5001). Transcription can be local (Whisper) or cloud (Gemini / DashScope); nothing is shared unless you configure a cloud engine.

Verbatim


What it does

  • Multi-engine transcription — pick per job:

    EngineBackendNotes
    Whisperfaster-whisper (int8/CPU, VAD), auto-detect language; falls back to openai-whisperLocal · free · offline
    GeminiGemini 2.5 Pro, 15-min chunking, retry + safety-filter diagnosticsCloud · high quality
    DashScopeAliyun Paraformer-v2 (async REST)Cheap · Chinese-tuned · native timestamps
    PreciseGemini text + DashScope speaker diarization, merged by Gemini per 10-min windowSpeaker labels · slow
  • Batch transcription — drag-drop many files; each runs as an independent task with its own progress row.

  • Channel Pipeline — paste a video / playlist / channel URL → yt-dlp download → transcribe → per-episode AI analysis → cross-episode synthesis. Downloads and transcription overlap; a clickable episode grid shows per-video status with thumbnails.

  • AI layers — structured summaries (overview + timestamped sections) and auto-generated card metadata (title / one-liner / tags) for every transcript.

  • Library — full-text search across titles, filenames, and transcript text; engine filters; a rendered Markdown reading view for analysis documents.

  • Personal dashboard — bottom-left panel: total hours transcribed, characters produced, a cumulative-hours line chart, and the topics you follow (from tags).

  • Exports — TXT, SRT subtitles, and downloadable Markdown documents.


Architecture

Browser (SPA: tabs, SSE progress, self-drawn charts, client-side Markdown renderer)
   │  fetch / SSE
Flask (getAudio/app.py)
   ├─ ThreadPoolExecutor + per-engine Semaphores      ── transcription concurrency
   ├─ global download / analysis Semaphores            ── pipeline throttling (shared across all chains)
   ├─ SQLite (taskdb.py)                                ── task state + restart recovery
   └─ filesystem: results/<uuid>/{audio, transcript.json, summary.json, meta.json}
                  results/_chains/<id>/{chain.json, 分析_*.md, 总分析.md}

Design principles baked in

  • Global, not per-request, throttling. Download and analysis concurrency is capped by shared semaphores, so running 5 pipelines at once produces the same external load as running 1 — they queue, they don't compound.
  • Overlap I/O and compute. In a pipeline, each video starts transcribing the moment it finishes downloading (wall-clock ≈ the slower of the two, not their sum).
  • Persist everything. Tasks live in SQLite; a restart re-queues unfinished work and cleans orphaned uploads. (Chains keep their state in chain.json but do not yet auto-resume — see Roadmap.)
  • Safe by construction. task_id is UUID-validated before any filesystem join (path-traversal guard); document reads reject .. / non-.md names; optional token auth gates the whole app.

Setup

Requirements: Python 3.9+, ffmpeg and yt-dlp on PATH (Homebrew recommended on macOS).

brew install ffmpeg yt-dlp          # yt-dlp MUST be the system binary, not a pip package
python3 -m venv venv && source venv/bin/activate
pip install -r getAudio/requirements.txt

Create getAudio/.env:

GEMINI_API_KEY=...          # for Gemini transcription, summaries, and pipeline analysis
DASHSCOPE_API_KEY=...       # for the Aliyun (DashScope) engine
# Optional:
# GETAUDIO_TOKEN=...             # enable token auth (blank = open, local use)
# YTDLP_COOKIES_BROWSER=chrome   # borrow browser cookies (fixes Bilibili 412, raises YouTube limits)
# WHISPER_LANGUAGE=zh            # force a language for local Whisper (blank = auto-detect)

Run:

bash getAudio/run.sh          # → http://localhost:5001

Or self-host with Docker

Everything (Python, ffmpeg, yt-dlp) is baked into the image — no local setup.

cp .env.example .env          # fill in keys + a GETAUDIO_TOKEN
docker compose up -d          # → http://localhost:5001

Transcripts and audio persist in ./data/. The image runs the app with the debugger off and binds 0.0.0.0. Notes:

  • Exposing to the internet? Set GETAUDIO_TOKEN in .env (visit /?token=... once), and put it behind HTTPS (a reverse proxy like Caddy/nginx). You pay all cloud-engine API costs.
  • The Channel Pipeline needs YouTube/Bilibili access. There's no browser in the container, so YTDLP_COOKIES_BROWSER is empty by default — for logged-in sources, mount a cookies.txt.
  • Local Whisper runs on CPU inside the container; heavy jobs are slow — prefer a cloud engine on a server.
  • The image is a few GB (PyTorch, for the openai-whisper fallback). Drop torch + openai-whisper from requirements.txt if you only use faster-whisper / cloud engines.

Configuration (getAudio/config.py)

SettingDefaultPurpose
ENGINE_CONCURRENCYwhisper 2, gemini 8, dashscope 9, precise 4Global per-engine transcription concurrency
CHAIN_DOWNLOAD_CONCURRENCY4Global cap on simultaneous downloads (all pipelines)
CHAIN_ANALYSIS_CONCURRENCY4Global cap on simultaneous analysis calls
YTDLP_LANGzh-CNyt-dlp metadata language (Chinese titles)
YTDLP_COOKIES_FROM_BROWSERchromeBrowser to borrow cookies from ('' disables)
WHISPER_MODEL_SIZEsmallfaster-whisper model
MAX_CONTENT_LENGTH500 MBUpload size limit

Key endpoints

Method / PathWhat
POST /upload, POST /upload_batchSubmit file(s) for transcription
GET /stream/<task_id>SSE progress stream
GET /api/history, /api/history/<id>, .../audioList / fetch / play saved transcripts
GET /api/search?q=Full-text search
POST /api/enrich_allBackfill AI titles/tags
POST /api/chainStart a Channel Pipeline
GET /api/chains, /api/chain/<id>, .../files, .../file?name=Pipeline status & documents
GET /api/statsAggregated dashboard data

Project layout

getAudio/
├─ app.py                  # Flask app: routes, workers, pipeline orchestration, auth, stats
├─ config.py               # config + concurrency/throttle params
├─ taskdb.py               # SQLite task persistence + restart recovery
├─ downloader.py           # yt-dlp probe/download (cookies, zh-CN titles, thumbnails)
├─ analyze.py              # per-episode analysis + cross-episode synthesis
├─ summarize.py            # AI content summary (Gemini / Qwen)
├─ enrich.py               # AI card metadata (title / one-liner / tags)
├─ transcribe_whisper.py   # local engine (faster-whisper → openai-whisper fallback)
├─ transcribe_gemini.py    # Gemini engine
├─ transcribe_dashscope.py # Aliyun Paraformer engine
├─ transcribe_precise.py   # Gemini text + DashScope diarization merge
├─ templates/index.html    # single-page UI
├─ static/{app.js, style.css}
├─ results/                # per-task outputs (+ _chains/ for pipelines) — git-ignored
└─ requirements.txt

Full source is also concatenated into getAudio/完整源代码_Verbatim.md.


Design language

The UI follows an Anthropic/Claude-inspired system: warm ivory ground (#FAF9F5), Claude coral accent (#D97757), serif display titles over small Apple-system body text. Status is carried by color + uppercase labels rather than emoji. All theming lives in CSS custom properties at the top of getAudio/static/style.css.


Known limitations / roadmap

  • Pipelines don't auto-resume after a server restart (transcripts do; chain state is saved but not re-run).
  • --cookies-from-browser requires the named browser installed locally.
  • Gemini may return empty on safety-filtered content; the pipeline marks that episode failed and continues.
  • Planned: route pipeline analysis to the Claude API for higher-quality reasoning (Gemini stays for cheap transcription); Gemini Batch API for 50%-cheaper, rate-limit-free analysis; pipeline auto-resume; SQLite FTS index for search at scale.

Personal project. Transcription/analysis of third-party content is for private study; respect the source platforms' terms and creators' rights.

Reviews

No reviews yet

Be the first to review this server!