← All work

Usher

Hackathon · MHacks 2026, University of Michigan · October 2026

Usher

A headset for people who cannot hear where a voice came from and cannot see it unless it is straight ahead. Say “Hello Jax”, and it turns the wearer toward you.

My part
Team lead: the idea and the research behind it, the hardware, the review and the test harness
Team
Jakhangir Tynshimov, Oliver Lai, Sajjad Akbari, Matthew Palmer
Built with
Raspberry Pi 5, reSpeaker XVF3800, vibration motors, SSD1306 OLED, Python, faster-whisper, ElevenLabs Scribe, Gemini, OpenCV, a printed headband
Event
MHacks 2026, University of Michigan, October 3⁠–⁠4, 2026
The Usher prototype on the build table: the white headset band wrapped in black tape and wired to the boards, the round microphone array on a stand above it
The prototype on the build night.

The trigger

“Hello Jax.”

Someone across the room says your name. Before you can join the conversation, you have to know where they are. The headset listens for a greeting followed by the wearer's name, works out where the voice came from, and tells the wearer which way to turn: by touch first, then on the display.

  1. The mic array hears it Four microphones on top of the head and an on-chip speech flag: the direction the voice came from, averaged over the phrase.
  2. A buzz on one temple A soft, continuous buzz on the side to turn toward. It fades as the wearer turns and stops when the speaker is in front.
  3. An arrow on the display LEFT, RIGHT, FRONT or BEHIND on a small OLED, confirming what the temple said.
  4. Catch me up The last thirty seconds of audio go out for transcription with speakers separated; what the caller said appears on the display.
  5. Live captions What people are saying, with an arrow toward whoever is talking now.
  6. Names and a summary An introduction turns an unknown speaker into a name; every twelve seconds a short summary of the conversation.
front behind
The wearer from above, front at the top. The dot is a voice; the temple nearer to it buzzes, and the display names the way. Move your pointer: it is the voice. Tap around the head: the tap is the voice.

Who it is for

Hearing loss, and a field of view that narrows

Usher syndrome is the most common genetic cause of combined deafness and blindness: hearing loss from birth, then retinitis pigmentosa narrows the field of view to a tunnel. Together they make one ordinary moment hard. You cannot hear which way a voice came from, and you cannot see the person unless they are straight in front of you. By the time you have found who is talking, you have missed what they said.

Every aid we found helps one of the two senses and fails the other: a cochlear implant gives speech but not direction, a visual sign language fails as the field narrows, a phone app answers in audio. Human interpreters do the thing we wanted the headset to do, by touch: who is talking, and where.

Three engines

Each for what it is best at

Offline

Whisper on the Pi

A rolling four-second window, transcribed every second. The wake rule: a greeting followed by the name. The name alone never fires, because people say names all the time.

Works with no internet: detection, the buzz and the arrow never need the cloud.

Live

ElevenLabs Scribe Realtime

Captions as people speak, and a faster second ear for the wake phrase: the same rule and a shared cooldown, so one phrase buzzes once.

Realtime cannot tell speakers apart, so the mic array does: each direction is a person.

By voice

ElevenLabs batch + Gemini

The catch-up and the judges' transcript: batch Scribe separates speakers by voice; Gemini 3.5 Flash-Lite names them from introductions and writes the summary.

Never on the buzz path. If the cloud fails, the wearer still gets turned toward the voice.

One rule shaped the whole build: the “Hello Jax” path runs on the Pi and works offline. Everything that needs the cloud runs behind it and can fail without touching the buzz.

What the judges saw

Both sides at once

A phone app is little use to someone with tunnel vision and hearing loss, so the web app is for the table, not the wearer. On the left, the headset camera with a tunnel-vision overlay the judge can tighten: what they see. On the right, a transcript with speakers told apart by voice and a running summary: what Usher gives them.

The judge web app in its demo mode: a camera view with a tunnel-vision circle on the left, a voice-identified transcript and a conversation summary on the right
The judge app, in its scripted demo: no headset needed.
The build desk from above: the round microphone array board, microcontroller and breadboards wired together, a laptop and a mouse among the cables
The other side of the table: the hardware, mid-build.
A teammate at a curved monitor showing the judge app: the tunnel-vision circle on the left and the transcript on the right
The tunnel-vision overlay on the big screen, late in the night.

Measured

Numbers from the table

Measured on the hardware during the night, with the mic on the printed band after calibration. Not a benchmark: one room, one team, and a hotspot.

“Hello Jax” to the arrow
1–2 s (Whisper on the Pi 5)
Direction error after calibration
3–11°, about 20 readings a second
Catch-up, 30 s of audio
0.6–2.4 s round trip
Batch transcript, 20 s
1.3–1.5 s, two speakers kept apart
Names and summary
every 12 s
Display
3 lines of about 20 characters, legible at 12 px

Troubles we hit

And what held

  1. Whisper heard “Jax” as “jacks” or “jack”.

    Match by sound (Metaphone), and accept “jack” only right after a greeting.

  2. The mic dropped off USB in the first full run and the captions died silently.

    An audio watchdog: two seconds of silence from the device, and it re-scans and reopens. Motors moved off the 5 V rail.

  3. Two sources renamed the same person back and forth.

    Local rules may only name the unnamed; only Gemini may change a name.

  4. Four people in a busy room became sixteen voices.

    A reference clip for every known voice in each request, 1.5 s of clear speech before a new one, quiet talkers ignored, a cap with a nearest-direction fallback.

  5. The first buzz pattern was too strong on the temples.

    A soft continuous buzz at 10–20% power that gets softer as the wearer turns, instead of counted pulses.

Twenty-four hours, in frames

A laptop showing the first version of the web app, 'Stay in the conversation', with Jakhangir in the camera feed
The first page, Saturday: the camera on me.
A teammate kneeling at an Elegoo 3D printer, printing the headband
The headband, printed overnight.
A breadboard taped to the printed headband, jumper wires fanning out of it
The breadboard rides on the band.
The prototype on the table: the Pi beside it, the OLED hanging on its wires, two face sketches on the whiteboard behind
The OLED on its wires; the whiteboard behind has the first sketches.
Jakhangir and two teammates lying on the floor between the chairs at night
Three in the morning.
A teammate seated wearing the headset, wires running to a laptop; Jakhangir standing beside him
A first fitting, Sunday morning.
At the judging table. 0:12
Fitting and explaining, Sunday morning. Sound · 0:40
The four teammates behind the MHACKS marquee letters in the judging hall
The four of us, after judging.

My part

What I did

I led the team. The idea was mine, and the research behind it: before the hackathon I read what people with Usher syndrome actually struggle with, what aids exist and where each one fails, and picked this from three candidates on the morning call.

The hardware was mine to bring and to wire: the reSpeaker array and the Pi, and the motor modules from TouchPoint, my earlier glove. The sound-direction and haptics approach carries over from SPOOT and TouchPoint.

Through the night I reviewed the code we had: sixteen findings with patches, an offline test harness of 144 checks that runs anywhere in three seconds, and the runbook we followed to the judging table. Oliver Lai wrote most of the program on the Pi and tested it on the device; Sajjad Akbari built the web page and the camera server; Matthew Palmer modelled and printed the headband and the mic tray.