miss-quote

Writes down what happened in the voice channel, and then tells you about it. When the bot leaves it summarizes the evening, and next time you can ask — out loud — "what happened last session", and a bard recounts the night. Built for a D&D table. It gets up to other shenanigans too.

Python 3.12 discord.py 2.4+ Wyoming ASR + TTS OpenAI-compatible LLM Silero VAD CPU-only container

Fair warning

Made with vibes, not love

This whole thing was vibecoded, it doesn't deserve proper effort, what it does isn't worth it. It was written as a joke, out of pure laziness, and yet somehow manages to do its job anyway. Anyone is welcome to it, but it comes with as much guarantee as effort that went into it: none.

The shape of it

Four things, in order

Audio handling is local and serial. Transcription is remote and parallel. Everything downstream is a tool that reads the stream.

Listen

Discord voice arrives as 48 kHz frames. soxr resamples and Silero VAD decides which of them are speech.

Transcribe

Each utterance is one connection to a Wyoming ASR server, dispatched at the speech-to-silence edge and bounded by a semaphore.

Write down

One JSONL file per session, per speaker, per channel — appended and flushed as produced, on a schedule you set.

Answer

Tools read the stream and reply through a Wyoming synthesizer, in the channel it was said in.

Cost

The GPU is somebody else's problem

This container needs no GPU. The Wyoming ASR and TTS servers it talks to do, and so does the endpoint behind a summary — budget for one, and let this share it. What that buys is a process cheap enough to sit in a voice channel all evening, scheduled anywhere, while the expensive hardware stays behind a socket several things can use.

~4.9 ms CPU per speaker per second of audio — about 0.5% of one core
~70 ms Round trip for a single utterance against a GPU-backed ASR
No database A config file, four volumes, and outbound calls — nothing to open a port for
1 process One event loop, one replica, no multiprocessing layer

Tools

What it does with what it heard

A tool reads a server's transcripts and does something with them. Each is opted into per server, and each one absent is reported at startup rather than left to be noticed.

🎬 quotes

Answers the channel with the film line it just walked into — and then asks where it came from. Name the title in time and you are paid a credit.

Configure quotes →

📝 summary

Turns a sealed session into an account of it, posts it to a text channel, and reads it back out loud when somebody asks what happened last time.

Configure summary →

📊 scoreboard

Keeps a running balance per person, writes it to disk, and publishes the standings under the name of whatever voice channel the bot is in.

Configure scoreboard →

🔊 tts

The only thing that plays anything. Owns the rendered-speech cache, the chime library, the volume, and the voice connection.

Configure tts →

🚨 verbal-morality

The Verbal Morality Bot, after Demolition Man. Fines a speaker out loud for the wrong word, and hands the fine to the scoreboard.

Configure verbal-morality →

🧩 Write your own

Three moments — an utterance, a sealed session, and a loop of its own. Declare what you need, reach your neighbours by class, and register the name.

The tool contract →

Quick start

Point it at a token and two servers

Everything a deployment points at is an environment variable. Everything about how it behaves is config.yaml.

docker run
# The token, the ASR, and the file that says which servers to join.
docker run -d --name miss-quote \
  -e DISCORD_TOKEN=... \
  -e WYOMING_HOST=asr.internal \
  -e TTS_HOST=tts.internal \
  -v ./config.yaml:/config/config.yaml:ro \
  -v miss-quote-transcripts:/transcripts \
  -v miss-quote-speech:/speech \
  ghcr.io/beer-wars/miss-quote:latest

config.yaml is not optional and is not shown here — it says which servers to join, and a server missing from it is never joined. The repository ships one commented in full to start from.

Full installation guide →

Read on

Everything else is documented

How the pipeline is split, what a transcript looks like on disk, every setting in the file, and every variable a deployment points at.