Real-Time Local Intelligence Engine

Real-time intelligence.
No cloud required.

Audrix captures live audio and on-screen visuals, transcribes with noise suppression and crosstalk detection, and runs OCR and Vision locally. Transcription runs 100% offline. Live AI analysis can run locally or via your preferred cloud provider, with 15 providers including OpenAI, Anthropic, StepFun, Groq, OpenRouter, and local Ollama and LM Studio.

Get Audrix Free See all features โ†“
Free to use 100% Offline
Audrix main interface showing real-time transcription and system audio capture
Portable Low Latency Offline Vision

Audrix Engine: real-time transcription from microphone or system audio, with local OCR and Vision for instant screen analysis.

๐ŸŽค System Audio
๐Ÿ”‡ Noise Suppression
17+ Models
100% Offline
0ms Cloud Latency

How it works.

Three steps to real-time local intelligence.

1

Choose a backend

Use WhisperLive (11 models), Lemonade NPU acceleration, or any external transcription endpoint. Switch at runtime without reinstalling.

โ†’
2

Capture audio

Grab microphone, system audio, or both simultaneously. Audrix handles noise suppression and crosstalk detection automatically.

โ†’
3

Get transcripts

See text appear in real time. Select any screen region for OCR and Vision. Everything runs locally on your machine.

Trusted by the community.

Reviewed and recommended by publications across the software space.

Softpedia

"Audrix is a tool that lets you capture sound from your microphone, PC speakers or both at the same time without dealing with complicated configuration. The transcription engine is designed to reduce duplicate words and improve accuracy, even if someone is speaking continuously or multiple people speak at the same time."

Read full review โ†’

Sponsored

Your audio and screen, analyzed instantly.

Audrix transforms live audio and on-screen visuals into immediate, actionable insights. Transcription runs 100% offline. Live AI analysis can run locally or via your preferred cloud provider.

Audrix is a real-time local intelligence engine for Windows. It captures microphone or system output for crisp speech-to-text with advanced noise suppression and crosstalk detection, while simultaneously running live AI analysis on the audio stream as it happens. Pair this with local OCR and Vision capabilities to instantly read and analyze any section of your screen.

Built for low-latency transcription with configurable noise suppression, automatic crosstalk detection, and 17+ selectable models. Transcription runs 100% offline. Your audio never leaves your machine.

Runs from a USB stick. No install, no account, no cloud sync.

Audrix settings and backend configuration

Sponsored

Why Audrix? Because local STT should be easy.

Common frustrations and how Audrix solves them.

AI can't hear your system audio

Most STT tools only capture the microphone. You cannot transcribe YouTube videos, Zoom calls, or any audio playing through your speakers.

  • Loopback capture: Record any audio playing through your system
  • Dual-stream mode: Capture mic and system audio simultaneously
  • Device selection: Choose specific input devices by system ID

Transcription locks you to the cloud

Cloud STT services send your audio to remote servers. You pay per minute, lose privacy, and depend on internet connectivity.

  • Fully offline: All processing happens on your machine
  • Multiple backends: 17+ models via WhisperLive, Lemonade, External endpoints
  • Zero recurring cost: Use free local models or bring your own API key

Overlapping speech ruins transcripts

In meetings or calls with multiple speakers, transcription becomes messy. Background noise and overlapping speech corrupt results.

  • Noise gate: Configurable threshold filters out silence and background hum
  • Crosstalk suppression: Dual-stream mode distinguishes user vs remote speakers
  • Speaker labeling: Automatic "User" / "Remote" tags in transcripts

Screen text stays locked in images

You can search audio but not what you see. Important details in videos, demos, or documents remain inaccessible to text-based tools.

  • Local OCR: Extract text from any screen region in real time
  • Vision analysis: AI-powered understanding of on-screen content
  • Region selection: Click and drag to capture any portion of your display

Switching engines means reinstalling

Most STT tools lock you into a single model or service. Changing transcription providers requires reinstalling or reconfiguring everything.

  • Hot-swappable backends: Change engines at runtime via configuration
  • Unified adapter interface: One API, multiple providers
  • No code changes: Switch without touching your workflow

Core features for serious audio intelligence.

Everything you need for real-time transcription, screen analysis, and flexible backend support.

Real-Time Transcription

Audrix transcribes audio in real-time as it is captured. See text appear the moment words are spoken, perfect for live captions, meeting notes, or voice commands. Smart merging prevents duplicate text, and partial results give you faster feedback before finalization.

System Audio + Mic

Capture your microphone, any app's audio output, or both simultaneously with automatic speaker separation. Transcribe your voice, YouTube videos, Zoom calls, or any system audio.

Noise & Crosstalk

Configurable noise gate and auto-gain keep background hum and silence out of your transcripts. Dual-stream mode distinguishes the local user from remote speakers on calls.

Flexible Backends

17+ transcription models via WhisperLive (11 Whisper variants), Lemonade NPU acceleration, or any external WebSocket or HTTP transcription endpoint. Switch at runtime without reinstalling.

Local OCR & Vision

Extract text from any screen region in real time with local optical character recognition. AI-powered visual understanding of on-screen content, diagrams, and UI elements.

100% Offline Transcription

Transcription runs entirely on your machine. Your audio never leaves your device. Live AI analysis can use local providers or your preferred cloud service - your choice.

Model Probing

One-click capability test for transcription backends. Know if a model handles accents, noisy audio, or long-form speech before you rely on it. The only reliable way to use a model at full capacity, because provider standards differ.

Custom Prompts

Swap the AI's behavior for live transcript analysis, vision, and OCR by switching prompt profiles. Interviewer mode surfaces questions and answers. Researcher mode extracts topics and references. Sales mode highlights objections and next steps. Build your own roles or import presets. The same model adapts to the job.

Dual-Stream Mode

Capture microphone and system audio simultaneously. Built-in crosstalk suppression distinguishes the local user from remote participants on calls.

Configurable Pipeline

Adjust sample rates, buffer sizes, commit intervals, and silence thresholds. Fine-tune every parameter for your hardware and use case.

Audrix main interface
Real-time transcription with live audio waveform and text output.
Audrix settings
Configure backends, audio devices, noise thresholds, and more.
Audrix models
Switch between Whisper, Lemonade, and external transcription engines.

See what you can ask Audrix.

From transcription to screen analysis, Audrix handles real-time audio and visual intelligence.

Transcription

  • "Transcribe this meeting in real time with speaker labels"
  • "Capture system audio from YouTube and save the transcript"
  • "Enable noise suppression and crosstalk detection for this call"
  • "Switch to the Lemonade backend for NPU acceleration"

Audio Routing

  • "Capture both microphone and system audio simultaneously"
  • "Route Zoom audio through VB-Audio Cable for transcription"
  • "Set the noise gate threshold to 0.02 for this environment"
  • "Label speakers as User and Remote in dual-stream mode"

Crosstalk & Noise

  • "Enable crosstalk suppression for multi-speaker meetings"
  • "Adjust RMS and peak thresholds for voice activity detection"
  • "Filter out background hum from the air conditioning"
  • "Separate overlapping speech into distinct speaker tracks"

System Audio

  • "Capture audio from this video and transcribe it live"
  • "Listen to system output without muting my speakers"
  • "Configure VB-Audio Virtual Cable for loopback capture"
  • "Transcribe a Zoom call while keeping my mic active"

OCR & Vision

  • "Select this screen region and extract all text"
  • "Analyze this diagram and describe what it shows"
  • "Read the error message from this screenshot"
  • "Extract table data from this on-screen spreadsheet"

Sponsored

How Audrix compares.

The local advantage for speech-to-text and screen intelligence.

What you get Audrix Cloud STT Generic Offline STT
Runs fully offline & privateTranscription is 100% local. Analysis runs on your machine or via your chosen cloud provider.Cloud only. Your data on their servers.Local processing, limited features
System audio captureLoopback + mic, dual-stream modeMicrophone onlyRare or unsupported
Crosstalk suppressionDual-stream distinguishes speakersBasic noise removal onlyUnsupported
Local OCR & VisionBuilt-in screen analysisCloud-only vision APIsUnsupported
PrivacyAudio stays local. Transcript analysis only sent if you choose a cloud provider.Audio sent to remote serversLocal but no ecosystem
Cost$0, free community build$0.006 to $0.03 per minuteFree but limited
NPU accelerationAMD Ryzen AI via LemonadeServer-side GPUs onlyCPU-bound
PortableRuns from USB, no installApp install or web onlyApp install required

Sponsored

Built for everyone who listens.

Researchers, coders, journalists, meeting hosts, and anyone who values privacy and offline capability.

Researchers & academics

Transcribe interviews, lectures, and meetings with perfect accuracy and speaker separation.

Writers & journalists

Dictate articles, transcribe interviews, and capture audio notes with instant text output.

Meeting hosts

Run live captions for Zoom, Teams, or Discord calls. Dual-stream mode separates local and remote speakers.

Coders & programmers

Transcribe coding sessions, capture system audio from tutorials, and use OCR to read documentation or error messages.

Privacy-conscious users

100% offline transcription with optional local or cloud AI analysis. No accounts, no cloud sync, no tracking unless you choose a cloud provider.

Sponsored

Supported providers

Transcription Engines

WhisperLive (11 models), Lemonade NPU (dynamic discovery), External WebSocket

WhisperLive
Lemonade NPU
External WebSocket

Live Analysis Providers

Cloud

OpenAI
Anthropic
Groq
StepFun
OpenRouter
NVIDIA NIM
Mistral
Gemini
DeepSeek
XiaomiMiMo
CommandCode

Local / On-premise

Ollama
LM Studio
Local Lemonade

A taste of what's inside.

Fast local transcription with AMD Ryzen NPU acceleration, plus external endpoint support.

Lemonade (Local) Qwen3.6-35B-A3B-UD-IQ2_M AMD Ryzen NPU-accelerated quantized model via FastFlowLM ACTIVE
Whisper whisper-large-v3 OpenAI Whisper large model for high-accuracy transcription CONFIGURED
External ws://localhost:9000 Custom WebSocket transcription endpoint EXTERNAL

Audio Capture deep dive.

Low-latency audio capture with configurable pipeline for transcription and analysis.

Low-Latency Capture

Audio is captured and processed with minimal delay. Reliable streaming delivers stable audio to the transcription engine.

  • Minimal overhead: Direct audio capture with no intermediate processing
  • Stable streaming: Consistent delivery to the transcription backend
  • Configurable buffer: Adjust frames per buffer and commit intervals

Setting Up System Audio (VB-Audio Cable Routing)

To correctly capture system audio without losing your loudspeaker output, Windows must be configured to pass the virtual stream back to your hardware. This requires VB-Audio Virtual Cable, a free virtual audio driver.

  • Install VB-Audio Virtual Cable and reboot if prompted
  • Open Windows Sound settings and switch to the Recording tab
  • Select CABLE Output, open Properties, and navigate to the Listen tab
  • Check "Listen to this device" and select your physical Loudspeaker
  • In Audrix, select Cable Input (VB-Audio Virtual Cable) as your device

Configuration Options

Fine-tune the audio pipeline for your hardware and environment.

  • Sample rate: Default 16000 Hz for optimal transcription quality, adjustable
  • Frames per buffer: 11025 default (configurable for latency/CPU trade-off)
  • Commit interval: How often to send accumulated audio to transcriber
  • Overlap buffers: Continuity between chunks for better word boundary detection
  • Auto-gain: Peak normalization with configurable target (default 0.25)
  • Silence thresholds: Peak (0.015) and RMS (0.003) for voice activity detection

Vision & OCR.

Read and analyze any section of your screen with local OCR and Vision.

Beyond audio, Audrix includes local OCR and Vision capabilities. Select a region, capture it, and get text extraction or visual analysis immediately. No cloud uploads, no API keys, no waiting.

  • Local OCR: Extract text from any screen region in real time
  • Vision analysis: AI-powered understanding of on-screen content, diagrams, UI elements
  • Region selection: Click and drag to capture any portion of your display
  • 100% offline: Every OCR and Vision model runs entirely on your machine
Audrix Vision and OCR settings

Get started in minutes.

Download, extract, and run. No installation required.

Choose from 17+ Whisper models via WhisperLive, Lemonade NPU, or external endpoints. Add live AI analysis with 15 providers, local or cloud.

Download

Grab the latest portable executable from GitHub. Extract it anywhere, even a USB drive.

Choose a backend

Use Whisper, Lemonade NPU acceleration, or any external transcription endpoint. Switch at runtime.

Start transcribing

Select your audio device, configure noise thresholds, and watch real-time transcripts appear.

Need help?

Join the community, leave a review, or support the project.

Community

Join the Discord server to ask questions, share workflows, and get help from other Audrix users.

Join Discord

Feedback

Have a question, suggestion, or bug report? Open a discussion or issue on GitHub.

GitHub Feedback

Softpedia feedback

Leave feedback on Softpedia to let other users know what you think.

Softpedia Feedback

AlternativeTo

Leave feedback on AlternativeTo to help others discover Audrix.

AlternativeTo Feedback

Simple pricing.

Audrix is fully free today. A planned premium option is coming for users who want extended use and support.

Free

$0

Free community build with a 30-minute session cap. Bring your own models or use built-in backends. No account, no trial clock.

Studio & Enterprise

Coming soon

Optional paid tiers for extended session duration, priority support, and commercial licensing. All current features remain available in the free tier.

Not yet available, join Discord for updates.

Join Discord for updates

Sponsored

Discover more apps.

Other tools from Tetramatrix for productivity, gaming, and AI-powered workflows.

Sorana Personal AI Knowledge Workspace

Sorana

Your personal AI knowledge workspace. A second brain that remembers your projects, learns your style, and acts on your files. Runs fully offline on Windows.

Aicono AI desktop icon organizer

Aicono

AI Intelligent Desktop Icon Autopilot. Automatically organizes your cluttered Windows desktop using AI. Group icons intelligently, arrange them neatly.

TabNeuron AI spatial tab manager

TabNeuron

AI Spatial Tab Manager & Research Workspace. Maps browser tabs onto a 2D canvas, AI groups them by content, chat with any page or the live internet.

New Spaceship retro arcade game

New Spaceship

Retro Arcade 2D side-scroller bullet-hell shmup game.

RyzenZPilot AMD power management

RyzenZPilot

Powerful tool for managing AMD Ryzen processor power settings on Windows. Adjust CPU performance, power limits, and thermal configurations.

Audrix main interface

Audrix

Real-time speech-to-text engine. Capture microphone or system audio, transcribe instantly with noise suppression and crosstalk detection. Runs entirely offline on Windows.

Common questions.

Quick answers before you download.

Does Audrix cost anything?

Audrix is free to download and use. A free community build is available with a 30-minute session cap. Studio and Enterprise editions are available for extended use.

Does my data leave my machine?

Transcription audio stays 100% local. Transcript text is only sent to the cloud if you choose a cloud LLM provider for live analysis. You can run entirely offline with local providers like Ollama, LM Studio, or Lemonade.

Can I run it from a USB stick?

Yes. Audrix is a portable executable. Extract it anywhere and run it. No installation, no registry changes, no admin rights required.

What transcription backends are supported?

Audrix supports WhisperLive (11 Whisper models: tiny through large-v3, plus distil-small.en and Parakeet TDT 0.6B), Lemonade NPU acceleration with dynamic model discovery, and any external WebSocket or HTTP transcription endpoint. Backends are hot-swappable at runtime.

Does it support system audio capture?

Yes. Audrix captures microphone, system audio (loopback), or both simultaneously. Dual-stream mode distinguishes the local user from remote speakers on calls.

Sponsored

Real-time intelligence. No cloud required.

Download Audrix and get instant local speech-to-text, system audio capture, noise suppression, and OCR. Free, private, and built for Windows.

Get Audrix Free View on GitHub