Local real-time voice transcription app powered by faster-whisper (OpenAI Whisper via CTranslate2). Zero cloud dependency — all processing happens on-device. Three modes: dictation (inject text at cursor), transcription (dedicated window with timestamped history), and file (upload audio files for offline transcription). Optional LLM post-processing (summarize, to-do list, reformulate) via any OpenAI-compatible API (Ollama, LM Studio, etc.).
Two-layer architecture communicating over localhost:
Tauri v2 (Rust shell + React frontend)
↕ HTTP + WebSocket (localhost:8001)
Python Backend (FastAPI + faster-whisper)
- Frontend (
frontend/): React 19 + Vite + Tailwind CSS v4 + shadcn/ui. Served by Tauri webview. - Tauri (
src-tauri/): Rust shell — global shortcuts, text injection (enigo), overlay window, system tray, sidecar process management. - Backend (
backend/): Python 3.13 + FastAPI — audio capture (sounddevice), faster-whisper transcription, file transcription, WebSocket streaming, SQLite storage, ChromaDB semantic search, LLM post-processing (OpenAI-compatible), config loading. - Config:
config.yamlat project root, validated by Pydantic inbackend/src/config.py.
| Layer | Technology | Key details |
|---|---|---|
| Desktop shell | Tauri v2 | Rust, global shortcuts via tauri-plugin-global-shortcut |
| Frontend | React 19, Vite 7, TypeScript 5.9 | Path alias @/ → frontend/src/ |
| UI | shadcn/ui, Tailwind CSS v4, Radix UI, Lucide icons | |
| Backend | Python 3.13, FastAPI, uvicorn | Async, WebSocket-first |
| Audio | sounddevice (PortAudio wrapper) | 16kHz mono, 80ms chunks |
| Transcription | faster-whisper (CTranslate2) | Whisper models (tiny → large-v3 / large-v3-turbo), CUDA or CPU |
| Storage | SQLite via aiosqlite | Sessions + timestamped segments |
| Semantic search | ChromaDB + paraphrase-multilingual-MiniLM-L12-v2 | 384-dim multilingual embeddings, ONNX, local CPU |
| LLM processing | openai SDK (AsyncOpenAI) | OpenAI-compatible API (Ollama, LM Studio, etc.) |
| File upload | python-multipart | WAV, MP3, FLAC, OGG, M4A, WebM, WMA, AAC, Opus |
| Text injection | enigo 0.6 + arboard 3 (clipboard fallback) | Win32 SendInput |
| Package mgmt | npm (frontend), uv (backend), cargo (Rust) |
All commands from project root:
# Frontend dev server (Vite, port 5173)
npm run dev
# Backend (FastAPI, port 8001)
npm run backend
# equivalent to: cd backend && .venv\Scripts\python -m src.main
# Tauri dev (launches frontend + Rust shell)
npm run tauri dev
# Backend tests
npm run test:backend
# equivalent to: cd backend && .venv\Scripts\python -m pytest tests/ -v
# Frontend tests (vitest)
cd frontend && npm test
# or watch mode: cd frontend && npm run test:watch
# Frontend lint
cd frontend && npm run lint
# Tauri build (production)
npm run tauri buildcd backend
uv venv --python 3.13
uv pip install -e ".[dev]"- Venv:
backend/.venv - Build system: hatchling, packages =
["src"] - Entry point:
python -m src.main - Python version: 3.13 (requires-python >= 3.12)
- Key dependency:
faster-whisper>=1.1.0(Whisper model downloaded automatically on first run)
openwhisper/
├── CLAUDE.md # This file
├── config.yaml # User config (gitignored, use config.example.yaml)
├── config.example.yaml # Config template
├── package.json # Root convenience scripts
├── prd.md # Product requirements document
│
├── backend/ # Python backend
│ ├── pyproject.toml
│ └── src/
│ ├── main.py # FastAPI entry point
│ ├── config.py # Pydantic config loader (config.yaml)
│ ├── exceptions.py # Custom exception classes
│ ├── api/
│ │ ├── _helpers.py # Shared helpers (_get_repo, _session_to_dict)
│ │ ├── routes/ # REST endpoints (split by domain)
│ │ │ ├── __init__.py # Aggregated router
│ │ │ ├── health.py # GET /health
│ │ │ ├── config.py # GET/PUT /api/config (hot-reload)
│ │ │ ├── audio.py # GET /api/audio/devices
│ │ │ ├── sessions.py # CRUD /api/sessions
│ │ │ ├── search.py # GET /api/sessions/search (semantic + exact match)
│ │ │ ├── llm.py # POST /api/sessions/{id}/summarize, /api/llm/*
│ │ │ └── upload.py # POST /api/transcribe/file
│ │ ├── ws.py # WebSocket endpoints (audio stream, file transcription)
│ │ └── _file_transcription_state.py # Pending file upload state registry (REST→WS bridge)
│ ├── audio/
│ │ └── capture.py # Microphone capture (sounddevice)
│ ├── llm/
│ │ └── client.py # LLM client (OpenAI-compatible: summarize, rewrite, scenarios)
│ ├── search/
│ │ ├── embedding.py # Multilingual ONNX embedding function (paraphrase-multilingual-MiniLM-L12-v2)
│ │ ├── vector_store.py # ChromaDB singleton (index, search, delete)
│ │ ├── backfill.py # Backfill existing sessions into ChromaDB
│ │ └── stopwords.json # Multilingual stopwords for exact-match filtering (7 languages)
│ ├── storage/
│ │ ├── database.py # SQLite init + migrations (V0→V1)
│ │ └── repository.py # CRUD for sessions & segments
│ └── transcription/
│ ├── whisper_client.py # faster-whisper integration (WhisperClient + WhisperSession)
│ └── file_transcriber.py # File-based audio transcription (streaming segments)
│
├── frontend/ # React + Vite
│ ├── package.json
│ ├── vite.config.ts
│ └── src/
│ ├── main.tsx # React entry point
│ ├── App.tsx # Router (react-router-dom v7)
│ ├── components/ # UI components
│ │ ├── Layout.tsx
│ │ ├── TranscriptionView.tsx
│ │ ├── StatusIndicator.tsx
│ │ ├── LanguageSelector.tsx
│ │ ├── BackendStatusBanner.tsx # Backend health check banner
│ │ ├── DeleteSessionDialog.tsx # AlertDialog for session deletion
│ │ ├── SessionSearchBar.tsx # Search + filter bar (semantic + metadata)
│ │ ├── ScenarioCards.tsx # LLM scenario processing buttons (summarize, todo, reformulate)
│ │ ├── ScenarioResult.tsx # LLM scenario result display (copy, dismiss, markdown)
│ │ ├── MicTest.tsx # Microphone test component
│ │ ├── LogoMark.tsx # Inline SVG logo component
│ │ ├── ThemeProvider.tsx # next-themes provider
│ │ ├── ThemeToggle.tsx # Dark/light mode toggle
│ │ ├── settings/ # Settings page section components
│ │ │ ├── RestartBadge.tsx
│ │ │ ├── SettingsGeneralSection.tsx
│ │ │ ├── SettingsAudioSection.tsx
│ │ │ ├── SettingsOverlaySection.tsx
│ │ │ ├── SettingsTranscriptionSection.tsx
│ │ │ ├── SettingsSearchSection.tsx
│ │ │ └── SettingsAdvancedSection.tsx
│ │ └── ui/ # shadcn/ui primitives (alert-dialog, sonner, ...)
│ ├── hooks/
│ │ ├── useWebSocket.ts
│ │ ├── useTranscription.ts
│ │ ├── useDictation.ts
│ │ ├── useFileTranscription.ts # File upload transcription workflow hook
│ │ ├── useSettings.ts
│ │ ├── useBackendHealth.ts
│ │ └── useTauriShortcuts.ts
│ ├── lib/
│ │ ├── api.ts # REST client to backend (sessions, search, config, LLM, file upload)
│ │ ├── tauri.ts # Tauri IPC bridge
│ │ ├── constants.ts
│ │ └── utils.ts # cn() helper (clsx + tailwind-merge)
│ └── pages/
│ ├── TranscriptionPage.tsx
│ ├── FileUploadPage.tsx # Audio file upload + transcription page
│ ├── SessionListPage.tsx # Session list with search/filter bar
│ ├── SessionDetailPage.tsx
│ ├── SettingsPage.tsx
│ └── OverlayPage.tsx
│
└── src-tauri/ # Tauri v2 (Rust)
├── Cargo.toml
├── tauri.conf.json
├── .cargo/config.toml # MSVC linker fix (see Build Environment)
├── capabilities/
│ └── default.json # Permissions: shortcuts, tray
└── src/
├── main.rs # Entry point
├── lib.rs # Tauri setup, plugin registration
├── shortcuts.rs # Global shortcut registration
├── tray.rs # System tray setup
└── injection.rs # Text injection (enigo/SendInput)
- Language: All code, comments, commit messages, and documentation in English.
- Frontend imports: Use
@/path alias (e.g.,import { cn } from "@/lib/utils"). - UI components: Use shadcn/ui. Primitives live in
frontend/src/components/ui/. - Backend structure: Domain modules (
audio/,transcription/,storage/,search/,llm/,api/), each with__init__.py. - Config access: Always through Pydantic models in
backend/src/config.py, never raw YAML parsing. - Tauri v2:
app.titledoes NOT exist — window title goes inapp.windows[].titleonly. - Tauri windows: Two windows defined —
main(900x700, resizable) andoverlay(100x36, transparent, always-on-top, click-through). - Global shortcuts:
Ctrl+Shift+D(dictation),Ctrl+Shift+T(transcription). - WebSocket-first: Audio streaming and transcription use WebSocket, not REST.
- Session modes:
'dictation','transcription','file'(audio file upload). - Frontend routes:
/(transcription),/sessions(list),/sessions/:id(detail),/upload(file upload),/settings,/overlay.
Critical: Git's link.exe shadows the MSVC linker. Fixed in src-tauri/.cargo/config.toml which explicitly sets:
- MSVC linker path:
C:\Program Files\Microsoft Visual Studio\2022\Community\VC\Tools\MSVC\14.44.35207\bin\Hostx64\x64\link.exe - Windows SDK lib/include paths for version
10.0.26100.0
Tauri CLI: Must be installed via npm install -g @tauri-apps/cli@latest (not cargo install — fails with linker issue).
Prerequisites:
- Windows 10/11 (64-bit)
- MSVC 14.44.35207 (VS 2022 Community)
- Windows SDK 10.0.26100.0
- Node.js 20+, Python 3.13, Rust 1.75+
- GPU optional: NVIDIA with CUDA 12.x for accelerated transcription (faster-whisper works on CPU too)
- Library: faster-whisper >= 1.1.0 (CTranslate2-based Whisper implementation)
- Model sizes:
tiny,base,small,medium,large-v3,large-v3-turbo - Device:
auto(CUDA if available, else CPU),cuda, orcpu - Compute type:
auto,float16,int8,int8_float16 - VAD: Silero VAD filter to skip silent regions
- Buffer: Configurable audio buffer (1–10 seconds) before transcription
- Architecture:
WhisperClientmanages model lifecycle,WhisperSessionhandles per-session audio buffering and transcription streaming - Model download: Automatic on first use, cached in
~/.cache/huggingface/
Optional integration with any OpenAI-compatible LLM API for text processing after transcription.
- Client:
backend/src/llm/client.py— singletonAsyncOpenAIclient, init/close lifecycle - Scenarios:
summarize(2-4 sentence summary),todo_list(extract actionable items as markdown checkboxes),reformulate(clean up filler words, grammar, transcription artifacts) - API endpoints:
POST /api/sessions/{id}/summarize— generate/update session summaryPOST /api/llm/process— process text with a scenario ({"text", "scenario", "language"})POST /api/llm/rewrite— rewrite text with custom instruction ({"text", "instruction"})
- Auto-summarize: When
models.llm.auto_summarizeis true, sessions are summarized automatically on end - Config (
models.llminconfig.yaml):enabled: bool (defaultfalse) — master toggleapi_url: str (defaulthttp://localhost:11434/v1) — OpenAI-compatible endpoint (Ollama, LM Studio, etc.)api_key: str (defaultollama)model: str (defaultmistral:7b)temperature: float (0.0–2.0, default0.3)max_tokens: int (64–4096, default512)auto_summarize: bool (defaulttrue)
- Frontend:
ScenarioCards(3 color-coded buttons) +ScenarioResult(display with copy/dismiss), shown in TranscriptionPage and FileUploadPage - Languages: 13 supported (fr, en, es, pt, hi, de, nl, it, ar, ru, zh, ja, ko) — system prompts adapt to target language
- Graceful degradation: All LLM features disabled if
enabled: falseor API unavailable
Upload audio files for offline transcription (instead of live microphone capture).
- Supported formats: WAV, MP3, FLAC, OGG, M4A, WebM, WMA, AAC, Opus
- Max upload size: Configurable via
max_upload_size_mb(50–1024 MB, default500) - Three-step flow:
POST /api/transcribe/file— upload file, create DB session (mode=file), save temp file- State registry (
_file_transcription_state.py) bridges REST upload → WebSocket connection WS /ws/transcribe-file/{session_id}— stream transcription progress + segments, cleanup temp file
- WebSocket messages:
status,file_info(audio duration),transcript_delta,progress(percent),session_ended,error - Frontend:
/uploadroute →FileUploadPagewith drag-and-drop, progress bar, language selector, ScenarioCards integration - Hook:
useFileTranscription— manages upload workflow states (idle → uploading → transcribing → completed) - Backend:
file_transcriber.py— async generator using faster-whisper's native ffmpeg decoding, quality filtering (compression_ratio, avg_logprob thresholds)
- Phase 1 (Foundations): Completed — Tauri + React + FastAPI scaffold, config loading, audio capture, overlay, system tray, global shortcuts, dictation mode.
- Phase 2 (Transcription mode): Completed — WebSocket frontend<->backend, React UI, SQLite storage.
- Phase 3 (Dictation + overlay): Completed — Text injection, overlay window, dictation mode.
- Phase 4.1 (Robustness & Tests): Completed — Error handling, backend tests.
- Phase 4.2 (Settings page): Completed — GET/PUT /api/config, GET /api/audio/devices, SettingsPage with hot-reload.
- Phase 4.3 (Session UX + Search): Completed — AlertDialog delete, optimistic delete with animation, toast notifications, ChromaDB semantic search, metadata filters (language, mode, date, duration), search endpoint, auto-backfill.
- Phase 4.4 (UI redesign): Completed — Renamed to OpenWhisper, warm stone + amber palette, Plus Jakarta Sans + JetBrains Mono fonts, dark/light theme toggle, new logo.
- Phase 4.4b (Model migration): Completed — Replaced vLLM/Voxtral with faster-whisper. No external server needed.
- Phase 4.5 (LLM post-processing): Completed — OpenAI-compatible LLM integration (summarize, to-do list, reformulate), ScenarioCards/ScenarioResult UI, auto-summarize on session end.
- Phase 4.6 (File transcription): Completed — Audio file upload, drag-and-drop UI, streaming progress, file_transcriber backend, WebSocket progress channel.
- Phase 4.6b (Search improvements): Completed — Multilingual ONNX embeddings (paraphrase-multilingual-MiniLM-L12-v2), auto-migration from English model, search state persisted in URL params.
- Phase 4.6c (Search quality): Completed — Summary-first indexing (summaries preferred over raw text), configurable distance threshold, exact keyword match ranking with multilingual stopword filtering, relevance scores (sqrt cosine) with UI indicators.
- Phase 4.6d (Modularisation): Completed — Split routes.py into routes/ package, split SettingsPage.tsx into section components.
- Phase 4.7 (Packaging & Release): Completed — setup.bat script, graceful config fallback (no crash if config.yaml missing), updated Pydantic defaults (model_size=auto, language=en, overlay=false), config.example.yaml aligned with defaults, .cargo/config.toml auto-generated per machine, uv run in package.json scripts.
- See
prd.mdfor full roadmap and feature backlog.
Two main tables:
sessions: id, mode ('dictation'|'transcription'|'file'), language, started_at, ended_at, duration_s, summary, filename (V1 migration)segments: id, session_id (FK), text, start_ms, end_ms, confidence
Migrations tracked via PRAGMA user_version (current: V1 — added filename column).
- Storage:
./data/chroma/directory (sibling to SQLitesessions.db) - Collection:
sessions— one document per session (summary preferred, full text as fallback) - Metadata per document: session_id, language, mode, duration_s, started_at
- Embedding model:
paraphrase-multilingual-MiniLM-L12-v2(384-dim, 50+ languages, ONNX, downloaded to~/.cache/chroma/onnx_models/on first use) - Embedding function: Custom
MultilingualEmbeddingFunctioninbackend/src/search/embedding.py(ONNX + tokenizers, no PyTorch) - Indexing strategy: Summary-preferred — indexes LLM summary when available (concise, topic-focused), falls back to full transcript text. Re-indexes automatically after auto-summarize or manual summarize.
- Search ranking: Two-tier ranking: (1) exact keyword matches first (all non-stopword query words found in document), (2) semantic-only matches second. Both tiers sorted by cosine similarity.
- Exact match logic: Accent-insensitive + case-insensitive keyword matching with multilingual stopword filtering (
backend/src/search/stopwords.json— fr, en, es, pt, de, it, nl). - Distance threshold: Configurable max cosine distance (
search.distance_threshold, default 1.0). Results beyond threshold are filtered out. - Relevance scores:
sqrt(max(0, 1 - cosine_distance))— displayed as percentage in UI with color-coded progress bar (green/amber/red) + "Exact" badge. - Config:
search.embedding_modelandsearch.distance_thresholdinconfig.yaml(configurable, auto-migration on model or strategy change) - Indexing: Automatic on session end (in
ws.py), re-indexed after summarization, deleted on session removal (inroutes/sessions.py) - Backfill: Auto-indexes existing sessions on first startup if ChromaDB collection is empty, embedding model changed, or indexing strategy changed
- API:
GET /api/sessions/search?q=...&language=...&mode=...&date_from=...&date_to=...&duration_min=...&duration_max=... - Graceful degradation: If ChromaDB init fails, search falls back to SQL-only filtering
| Service | Port | Protocol |
|---|---|---|
| Vite dev server | 5173 | HTTP |
| FastAPI backend | 8001 | HTTP + WebSocket |