
The evolution of engaging with AI has mostly been:
chat.
Type something. Wait. Get something back. Repeat.
And honestly, it’s been incredible.
But after using these systems constantly for the last few years, I increasingly think chat is just the first interface layer.
Now most people seem to be converging on something else:
- voice works great for input
- visuals work great for output
Talking is often faster than typing.
And for productivity workflows, the thing we actually want back from AI usually isn’t audio.
It’s visual artifacts:
- code
- docs
- spreadsheets
- dashboards
- slides
- interfaces
- edits

I think we’ll absolutely see more video too, but I suspect that dimension is mostly about humanizing AI.
For actual work, the visual artifacts matter more.
So then the question becomes:
what’s the interface after chat?
The core interaction loop with AI is actually quite simple.
We look at something. We ask AI to do something. AI does something. We either:
- confirm it
- reject it
- refine it
- interrupt it
- or continue the conversation
Most of the interaction is basically:
yes. no. other.
And increasingly our computers feel like weird slot machines.
We bounce between windows and tabs with:
- variable response times
- intermittent rewards
- occasional brilliance
- occasional nonsense

So I started wondering:
what if the system simply knew what we were looking at?
My first instinct was some kind of face-worn device.
Almost like a Duck Hunt interface for AI.
Not literally.
But the interaction model felt similar: point attention somewhere → signal intent.
Then I realized webcam-based eye tracking is already pretty good.
Maybe we don’t even need something on our face.
The computer already knows:
- cursor position
- active windows
- scroll behavior
- dwell time
- screen regions
Add lightweight gaze estimation and it’s probably good enough.
So then I moved to the gesturing side.
At first I thought maybe the right answer was a tiny 3-button mouse.
Part of me still loves that idea.
Maybe we finally bring back Doug Engelbart’s input device after all.
But eventually a haptic ring started making the most sense.
Minimal. Always there. No visible buttons.
Just:
- tap → yes
- stronger squeeze → no
- squeeze for 0.5s → start speaking
Or honestly maybe speech sensitivity gets good enough that even the hold gesture becomes optional.
The more I thought about it, the more it stopped feeling like a gadget and started feeling like punctuation for thought.
So I started using ChatGPT to render out concepts as I refined the idea.
Sketches. Interaction models. Industrial design directions. High fidelity concept renders.
And somewhere during the process it crossed from:
“interesting thought experiment”
to:
“wait, I would actually use this today.”
I genuinely think something in this direction eventually become mainstream.
Not because it feels futuristic.
But because it feels lower friction.
And historically, the interfaces that win are usually the ones that disappear.
Also, my Apple Vision Pro already does this by looking at my eyes and my fingers. This concept is just a much cheaper, lighter, lower friction alternative.
