A new interaction model for real-time, on-display language translation on Meta's first AI glasses with high resolution in-lens display — patent filed, launched at Meta Connect 2025.
Lead Product Designer
(sole designer)
1 PM
4 engineers
2 content designers
1 UX researcher
Interaction Model team
User Education team
Dec 2024 - Oct 2025
Visual Translation is a real-time written language translation feature on Meta's first display AI glasses. You can look at foreign text in the world — a menu, a street sign, a poster — and read the translation on the glasses' in-lens display directly. With the multimodal AI silent-entry strategy I co-authored, you won't even have to say "Hey Meta" out loud to trigger the visual translation experience.
・・・・・・・・・・・・
It's a feature with no precedent on the form factor. Behind the scenes, we had tackled a combination of challenges, including 600 x 600px in-lens display size, text legibility on the small display, and variations and complexities in orientation, arrangement, and length of the real-world texts. The latter is where my engineer and I iterated A LOT on so that the system can behave naturally as how human reads. As one of the solutions, we also invented smart auto-zoom snapping to ensure legibility, ultimately making interactions smoother and smarter with minimum user input.
・・・・・・・・・・・・
I was the sole designer on Visual Translation and other multimodal AI features utilizing photo and OCR (optical character recognition) input from 2024 through its Day-0 launch on Meta Ray-Ban Display and Neural Band in 2025. I designed the interaction model from scratch with Jiaqian Wu and other engineers, refining it together for easy consumption of real-world text and its translation. The feature shipped as one of the launch experiences for multimodal AI on the display glasses, announced at Meta Connect 2025 keynote, and has a patent filed for its interaction model with me as the first inventor.
Visual Translation was one of the rock multimodal AI experiences for the new form factor, display AI glasses that Meta had been brewing for years. It aimed to add value proposition to the new consumer wearable and find its product market fit by allowing users to take actions on what they see, translating on the go, all without the need to reach for their phone, hence staying connected with the world and the people around them.
I categorized translation scenarios around 2 use cases:
As sole designer, I owned the interaction model, on-display rendering & navigation, gesture vocabulary, and the design POV for silent trigger for multimodal AI experiences. I partnered with the Interaction Model and User Education teams to shape and define the platform-wide gesture patterns and contextual gesture tutorials.

Visual Translation is one of those experiences that had to be iterated heavily based on the feedback from how users interact with it. In the initial stages, we learned the following challenges from user research studies and internal employee testings (dogfooding):
Each problem on its own felt like a fix. Together, they were a trust problem. If the user can't discover, can't navigate, and can't read the translation, the AI feels broken regardless of whether the model is right.
I anchored the work on 3 principles to guide the redesign directions:
These principles funneled into a clean framing:
"Optimizing for legibility of short texts and understanding of long texts."
2 scenarios, 2 different priorities, 1 consistent interaction language.
From there, I broke down the experience into 3 questions that guided the focus of design decisions:

The biggest design decision was the gesture vocabulary. The existing model — Captouch zoom + Head IMU pan — was producing too much friction for the user and too much instability for the visual experience. I proposed a new sets of interaction model that work together and complement each other, innovating on how users move through and consume real-world text and its translation without direct-manipulation touch screen while ensuring legibility:
Reducing manual manipulation and increasing automation through the smart text selection was the answer, both visually for the user and technically for the system. Well for the most part, that's why we still enable free panning and zooming because nothing worse than user feeling stuck in the experience.

There were a lot of detailed considerations and guidelines needed in place for engineering to implement a seamless smart panning + auto-zoom snapping experience, from defining the max zoom level to requesting suitable image resolution. Some of these tradeoffs were real, including the tension between legibility and latency, and how we could mask the latency with transition animation. There were also significant iteration and calculation that went into having the system understand the orientation and reading order like a human, so that when user swipes, the next selection aligns with what user intends and expects. We also drilled into the time threshold for the system to know if you intend on landing on a specific word, so the auto zoom animation is smooth and blends well with selection.

However, the text blocks are constained by the OCR (optical character recognition) grouping, plus the 600x600px display constraint, it is obviously not conducive for long-form text consumption, so I designed a solution for users to read paragraphs on the in-lens display if necessary.
I redesigned this whole launch experience based on executive, UXR, internal employee testing and design system feedback, refining and optimizing the interaction logic in lock-step with Jiaqian.
This proposal didn't just live inside Visual Translation. Working with the Interaction Model team — in collaboration with Alex Gerrese — I shaped the broader proposal to unify wrist-roll zoom and pan gestures across other features, such as Map and Gallery. The interaction language I designed for translation became platform language.
Visual Translation along with all other multimodal AI features I standardized the interaction model pattern across were all launched with the announcement of Meta Ray-Ban Display with Neural Band at Meta Connect 2025 by Mark Zuckerberg.
Jiaqian and I also filed the patent for real-world text consumption interaction model on Augmented Reality device on the same day!





After Visual Translation is shipped alongside all the other multimodal AI experiences I designed for Meta Ray-Ban Display, the first round of post-launch user research and dogfooding surfaced 3 issues:
I designed two solutions, in partnership with the Interaction model and User Education team:
The Day-90 release closed loops that Day-0 had left open. It also gave me a sharper picture of the gesture-discoverability problem on glasses generally — a problem I'd seen coming at the launch but hadn't fully solved.
Shipped on Meta's first AI glasses with high-resolution full-color in-lens display with whole new interaction model on the Neural Band, including the Day-0 multimodal AI launch and the Day-90 refinement release.
Patent filed for the interaction model for real-world text consumption on Augmented Reality (AR) devices by Meta Connect 2025, with myself recognized as the first inventor — my first patent.
Meta Ray-Ban Display glasses and Neural Band named to TIME's Best Inventions of 2025 and won UX Design Award 2026.
Visual Translation launched at Meta Connect 2025; other multimodal AI features I shipped on glasses were featured in both 2024 & 2025 Mark Zuckerbuerg's keynotes — covered by The Verge, The Verge hands-on, UploadVR, CNET, and The Verge accessibility column.
I would build a tighter feedback loop with the ML model, OCR, and Camera teams. Some of my design decisions were band-aids for upstream model behavior, OCR, and camera limitation. Earlier design–ML collaboration would have let me shape and influence the model's UX-relevant behavior directly, instead of designing around it downstream or heavily relying on my engineers to communicate with them.