RESONÆ

Voice AI tuned the way instruments are tuned — by ear, inside a cathedral built of standing waves.

Every arch in this room is an oscilloscope trace. Scroll, and the building answers your speed. A fictional lab · a real argument
about how synthetic voices should behave

Built for the three acts of speech.

Most audio stacks treat talking, hearing, and trusting as separate products. Resonae rehearses them in one room, because a voice you can fake, miss, or doubt is a voice you cannot use.

Act I · Sing

Synthesis with breath in it

Resonae performs a script rather than reading it. Phrasing, hesitation, and warmth are direction you write — not artifacts you fight.

sung in sub + bass
20–250 Hz
Act II · Hear

Comprehension that keeps the room

Words, speakers, and feeling, lifted out of live audio together — transcripts that remember who was talking and how they meant it.

heard in low mid + mid
250 Hz – 2 kHz
Act III · Sign

Provenance you can play back

Every generated second carries an inaudible signature. Run it through the public verifier and the watermark rings back true.

signed in presence + brilliance
2–20 kHz

Tuning is not decoration here; it is the work. One bench session from the fiction, logged the way it would read — a single voice walked down the same six bands the HUD in the corner sweeps.

# tuning transcript — bench session 041 · fiction, like every product here
# instrument: nave · voice "verger" · consent on record in crypt

00:00.4  SUB         31 Hz    blower rumble under the take     −29 dB  → cut to the floor
00:01.2  BASS        142 Hz   chest resonance of the voice     −7 dB   → held — the warmth stays
00:02.7  LOW MID     340 Hz   standing room mode, near wall    +4 dB   → nulled
00:04.1  MID         1.1 kHz  vowel honk off the hard floor    +3 dB   → eased flat
00:05.8  PRESENCE    3.4 kHz  consonant edge                   buried  → +1.5 dB, words land
00:07.3  BRILLIANCE  7.1 kHz  sibilance on "the nave is open"  +5 dB   → shaped back
00:08.9  signature   15 kHz   watermark laid across the take   absent  → in, inaudible
00:09.4  reply       capture to first voiced sample   412 ms  → 240 ms

verdict: TUNED — every band at rest · released to the rack

Six instruments in the rack.

Named for the rooms of the building they were tuned in. Each does one thing, does it end to end, and hands its output to the next without ceremony.

Nave

Live conversation engine

Speech to speech in one breath. Nave answers fast enough to be interrupted and is polite enough to stop when you do.

duplex · barge-in · 16 kHz in / 48 kHz out

Choir

Ensemble synthesis

One take, many voices. Cast a scene, hand Choir the script, and record the whole room at once — interplay included.

12 voices per take · shared room tone

Apse

Audio comprehension

The listening end of the building. Apse turns raw recordings into who said what, when, and in what mood — with receipts.

diarisation · sentiment · word timestamps

Transept

Cross-lingual voice

Your voice, crossing languages without leaving you. Transept keeps timbre and intent while the words change country.

9 languages in rehearsal · timbre held

Crypt

Provenance vault

Where the signatures live. Crypt keeps consent recordings and watermark keys, and answers one question: is this voice real?

consent ledger · public verifier · key escrow

Spire

On-device compiler

The whole cathedral folded to fit a pocket. Spire compresses a model until it runs on the phone in your hand — offline, private, yours.

int4 · 40 MB footprint · no network required

Principles before product.

Resonae is a fiction — a concept lab built to argue in public about how voice AI ought to behave. No customers, no wait-list, no funding round.

Honest framing: every product on this page is imagined. The principles are not. Take them; they are the point of the exercise.

Consent is the instrument

No voice is cloned without its owner on the record — literally. Crypt keeps the recording of the yes, next to the keys it protects.

Say when it is synthetic

Every output is watermarked, and every watermark is verifiable by anyone — not only by us. Provenance that needs our permission is not provenance.

Intimacy is not a growth channel

A voice in your ear is trust at close range. Nothing here pretends to be a person, and nothing here is tuned to keep you talking.

Local beats loud

If a model can run on your device, it should. Audio that never leaves the room can never leak from it.

The interface is a stage direction.

You do not configure a Resonae voice; you direct it. One endpoint per instrument, plain-language direction instead of markup, and the watermark on before you ask.

# one call, one performance
POST https://api.resonae.ai/v1/perform

{
  "instrument": "choir",
  "script": [
    { "voice": "verger",
      "line": "The nave is open.",
      "direction": "hushed, smiling" },
    { "voice": "cantor",
      "line": "Mind the standing waves.",
      "direction": "warm warning" }
  ],
  "output": { "format": "wav", "rate": 48000 }
}
  • Direction, not markup. "hushed, smiling" beats a pitch contour tag you will never tune by hand.
  • One endpoint per instrument. Perform, comprehend, translate, verify — four verbs, no console safari.
  • Watermark on by default. The verification endpoint is public and free to call. It stays that way.

Step into the nave.

The doors are fictional; the argument is not. Write to the lab, borrow the ideas, or take the whole build apart in the notes.

Arrival · all waves at rest · 20 kHz