Analysis · Alibaba
We Decoded Qwen's Chinese Slides: Inside Alibaba's Agent Smart Glasses
Alibaba's Qwen agent glasses were unveiled with dense Chinese engineering slides that nobody translated. We ran image recognition on all of them, so here is every menu, every diagram and every spec, in plain English.
Alibaba's Qwen unveiled its next-generation display smart glasses at the World Artificial Intelligence Conference in Shanghai, which opened on July 17, 2026, and the pitch is no longer "AI glasses" but agent glasses: a device that calls third-party Skills and Agents on demand, powered by Qwen, Alibaba's flagship multimodal model. The hardware sheet is aggressive - full-duplex audio, high-precision eye tracking, an always-on camera path that runs at 5 milliwatts, 3DoF spatial display, plus PPG and temperature sensors for continuous vitals.
Here is the part nobody did. The technical slides Qwen showed are entirely in Chinese, packed with block diagrams, mode labels and mock UI bubbles, and the English coverage so far has only repeated the press summary. We ran image recognition across every slide and transcribed the labels one by one. What comes out is far more specific than the announcement: you can read the wake-up hierarchy, the audio pipeline, the privacy gate, and even the fake notifications Qwen chose to render inside the lens.


What it implies: an eye-tracking module under 1.7 cubic millimetres drawing less than 100mW is the kind of number that decides whether a feature ships in eyewear or stays in a headset. Meta's Ray-Ban Display has no eye tracking at all, which is exactly why it needs the Neural Band on your wrist to move a cursor. Mentra, the open-source camp, ships no gaze sensor either and leans on voice plus phone touch. Google and Samsung have the harder version of this problem: Android XR already assumes gaze plus hand input on Project Moohan, a headset with power and volume to spare, but Galaxy-branded eyewear built with Gentle Monster and Warby Parker has neither. If Qwen really has gaze at this size and power, Alibaba gets a pointer that costs nothing in weight while its rivals keep paying for a wristband or a phone.
Start with the voice stack. The slide title reads 免唤醒连续对话的语音全双工架构, "full-duplex voice architecture for wake-word-free continuous conversation", with a subtitle promising 精准抗干扰、动态判停、实时意图接管: precise interference rejection, dynamic end-of-speech detection, and real-time intent takeover. The left half is labelled 流式 AudioLLM 预训练 (streaming AudioLLM pre-training), where an audio encoder feeds an adapter into the Qwen LLM alongside text tokens. The right half, 全双工后训练 (full-duplex post-training), is the interesting bit.
That post-training panel exposes a four-state classifier the press release never mentions. Reading the legend: <0> means 噪声或者静音, noise or silence. <1> means 有语音但无实际语义, speech present but no real semantics. <2> means 有语义但不完备,语音可打断信号, semantically meaningful but incomplete, an interruptible signal. <3> means 语音语义完备,判停信号, semantically complete speech, the stop signal. The mock inference shows the tokens 那 北京 明天 的 天气 呢 tagged <1>, <1>, <2>, <2>, <2>, <3>, with an annotation that two consecutive <2> tags trigger 打断 (interrupt) and a <3> followed by <0> triggers 判停 (end of turn). That is how Qwen gets overlapping conversation without a wake word, and how it claims a 90% cut in false triggers in noisy rooms.


What it implies: killing the wake word is the single most valuable interaction change on the table, and Qwen is attacking it at the token level rather than with a better trigger word. Every "Hey Meta" is a reminder that the glasses are a remote control, not a companion, and Meta's own live translation still works in turns. Google is the one rival with a genuinely comparable asset, since Gemini Live is already streaming and interruptible, but shipping that inside Samsung eyewear means fighting battery and thermal limits that a phone does not have. Mentra, being open, will likely wire in whatever streaming model is cheapest and let developers tune the barge-in themselves. The four-state classifier is also the honest answer to false triggers: you cannot claim always-listening glasses in a noisy street unless the model can tell background chatter from an unfinished sentence addressed to it.
The microphone chain gets its own slide, titled 端云协同的全链路拾音架构, an edge-cloud collaborative full-chain audio pickup architecture whose stated goal is 缓解端侧算力依赖 提升翻译 / 听记准确率: reduce reliance on on-device compute, improve translation and transcription accuracy. The 端侧链路 (edge chain) runs 5Mic 阵列 + 1VPU, then 回声消除 (echo cancellation), then 定向增强 (beamforming), then 编码 (encoding). Multi-channel audio then travels through the 手机 (phone) into the 云端链路 (cloud chain): 解码 (decode), 丢包补偿 (packet loss compensation), 增益控制 (gain control), and finally 大模型语音增强, large-model speech enhancement. The product shot is captioned 骨传导麦克风, bone-conduction microphone. Five mics plus a bone-conduction pickup is, as Qwen claims, the highest mic configuration currently shipping in this class.


What it implies: Alibaba is admitting that on-device compute is not enough and pushing speech enhancement into the cloud, which is a very different bet from Meta's. Meta keeps translation and captions as close to the device as it can because latency and offline behaviour are part of the pitch, and because sending raw room audio to a server is a regulatory fight in Europe it does not need. For Samsung and Google the same architecture is both easier and riskier: easier because Gemini already lives in the cloud, riskier because Galaxy glasses will be sold in the EU on day one and a cloud audio pipeline invites the data-protection questions Meta has spent two years absorbing. Mentra developers, meanwhile, get a template they can copy with any hosted model. Five microphones plus bone conduction is the real differentiator here, and it is a hardware cost nobody in the sub-300 dollar tier is paying.
The eye-tracking slide is where the numbers hide. The 高精度眼动追踪系统 panel lists four claims: 高精度 with 识别偏差 < 1°, a recognition error under one degree; 高鲁棒性 with 0-10万 Lux 全场景光照环境可用, usable from 0 to 100,000 lux; 低功耗 with 运行功耗 < 100mW; and 小体积 with 眼动模组 < 1.7mm³, an eye-tracking module under 1.7 cubic millimetres. A second slide splits the feature in two: 眼动交互 (gaze interaction) with the tagline 直觉化操控 视线即指令, intuitive control where the line of sight is the command, and 注视点识别 (gaze-point recognition), which promises 实现对复杂环境中单一目标的"即时锁定", instant lock-on to a single object in a cluttered scene.
That second panel contains the only true UI mockup in the deck, and it is a museum scene of blue-and-white porcelain. The user bubble says 这是什么?, "what is this?", and the glasses answer 你左前方的是..., "the one to your front left is...". Read together with the spec sheet, the intent is explicit: gaze replaces the touchpad, and the fixation point becomes a crop hint for the vision model, so "what is this" no longer means "describe the whole frame". Alibaba is also routing gaze hardware into payments, since the iris verification built with Ant Group rides on the same optical path.


What it implies: "what is this?" only works if the glasses know what "this" is, and gaze is the cheapest way to answer that. Today Meta answers the whole frame, which is why its visual replies are generic in a cluttered scene, and why the Neural Band exists as a substitute pointing device. Mentra has the same framing problem with no sensor to solve it. Google is the one that should be nervous, because gaze plus a multimodal model is precisely the Gemini demo everyone remembers, and Alibaba may put it on a face before Android XR eyewear reaches shelves. Add Ant Group iris verification on the same optical path and gaze stops being an input method: it becomes an identity and payment sensor, a business Meta cannot copy quickly and one Samsung would have to negotiate with banks region by region.
The always-on camera is not the main camera at all, and the slide makes that unusually clear. Under 端侧部署视觉识别模型 (on-device visual recognition model) it states 低功耗感知芯片 运行功耗 < 5mW. The flow labelled 分级视觉唤醒机制, a tiered visual wake-up mechanism, goes 高清摄像头 (HD camera) into 低功耗感知芯片 (low-power sensing chip, handling 控制/通信/OTA, image recognition, Jpeg/YUV encoding and RAW processing) into 本地服务决策系统, a local service decision system. From there it forks: AON 模式 leads to 各场景主动服务 (proactive services per scenario) entirely on device, while 高清模式 wakes the 高性能计算芯片. Only past that fork do you hit the row marked 图片 / 隐私授权 / 用户上传 - image, privacy authorisation, user upload - before anything reaches 视觉记忆千问大模型 (the Qwen visual-memory model) and 千问 AIGC 大模型. In other words, the consent gate sits between the phone and the cloud, not between the sensor and the chip.


What it implies: a 5mW perception chip watching the room all day is the feature that turns an assistant into an agent, and it is also the feature that will define the next privacy fight. Note where the consent gate sits on this slide: between the phone and the cloud, not between the sensor and the chip. Meta has deliberately not shipped an always-on camera path, and its capture LED is now part of its legal defence in a market where New York has already banned smart glasses in courthouses. Google and Samsung inherit that constraint multiplied by the EU AI Act and their own enterprise customers, so a continuously watching Galaxy pair is close to unshippable in the West without a very loud hardware indicator. Alibaba can ship it in China first, gather the behavioural data, and let its rivals argue about consent while it learns what proactive service actually feels like.
On display, Qwen is breaking with the head-locked panels that dominate current AI glasses. The 3D 空间显示 slide splits into 3D 显示, built on 双显双光机 (dual displays, dual optical engines), 双目视差渲染 (binocular parallax rendering) and 亚毫秒级硬同步 (sub-millisecond hardware sync), and 3DoF 显示, built on 实时精准头部追踪 (real-time head tracking), 双芯协同渲染 (dual-chip collaborative rendering) and 帧率高至 90FPS. Screens are anchored in the room rather than pinned to your nose, and 7.1 virtual surround audio tracks head pose so sound sources stay physically placed.
The mock cards in that render are worth reading too, because they telegraph the launch use cases: a calendar tile showing 7月 星期五 17, a weather tile reading 晴转多云 22°/29° with AQI 21 空气优, a notification 来自张三 下午4:00开会 会议室2 ("from Zhang San, meeting at 4pm, meeting room 2"), a chat thread 这个是设计稿 / 麻烦审核一下 / 没问题, and a large floating panel with the lyrics of Beyond's 《海阔天空》. Calendar, weather, work IM and lyrics: this is a productivity and everyday-companion device, not an AR gaming pitch.


What it implies: this is the clearest gap with Meta. Ray-Ban Display is monocular and head-locked, a notification surface pinned to your nose, while Qwen is describing two optical engines, binocular parallax and 3DoF head tracking at 90FPS, which anchors panels in the room. Viture and XREAL already sell 3DoF, but as tethered viewers, not as agent glasses. For Samsung the stakes are direct: Galaxy glasses with Gemini will be judged against this, and Android XR needs a spatial window manager for eyewear rather than a scaled-down headset UI. For Google it is a platform question, because whoever defines where a floating panel lives, how it follows you and what latency is acceptable, defines the app model. The mock cards tell you the target use cases too: calendar, weather, work chat and lyrics. Productivity and daily companion, not AR gaming.
The health slide, 用户体征实时监测系统, is the one with the most engineering detail per square centimetre. The 端侧 column chains AON signal acquisition into ET (eye tracking), PPG and IMU into 滤波消噪 (filtering and denoising) then 多维度特征提取 (multi-dimensional feature extraction), and NTC 环温补偿 (ambient temperature compensation) into 红外测温 (infrared temperature measurement). The App column handles 数据融合和统计 and 生理指标计算, passing HRV and heart rate up to the 云端 column, which runs 多模态生理压力算法 (multimodal physiological stress), 心率/血氧算法 (heart rate and blood oxygen) and 多模态发热风险算法 (fever risk). A banner underneath says Qwen built a 「主任医师级」大模型 with medical experts - a "chief physician grade" model - and the footnote warns 本功能提供的健康信息仅供日常参考,不构成医疗建议, health information is for daily reference only and is not medical advice.
The health mockup shows the interaction Qwen is really selling. The user asks 最近总觉得累,怎么回事呢?, "I keep feeling tired lately, what's going on?", and the glasses reply 今天久坐超过3小时,压力指数也比上周高20%,多户外活动有利于身心平衡: you have been sitting for over three hours today, your stress index is 20% higher than last week, more outdoor activity would help. Three tiles label the pillars: 生理压力 (physiological stress), 体征监测 (vitals monitoring) and 长期记忆 (long-term memory). That last one matters more than the sensors, because it is where a wearable stops answering questions and starts holding a profile of you.


What it implies: health is the flank where Meta is weakest and Samsung is strongest. Meta's sport line with Oakley counts activity and heart rate through partners, and Mentra has no sensor story at all, so a PPG plus temperature plus eye-fatigue pipeline in a display pair puts Alibaba in a category Meta has avoided for regulatory reasons. Samsung, on the other hand, already owns the watch, the ring and Samsung Health, so glasses that read HRV and fatigue would slot into an ecosystem that Alibaba has to build from scratch, and Google would supply the model that interprets it. The caveat is in Qwen's own footnote: this is daily reference, not medical advice, and a "chief-physician-grade" model is a marketing claim, not a clearance. Whoever tries to sell fatigue detection in Europe or the United States will need more than a slide.
Our take: read as a whole, the deck is a bet that the bottleneck in AI glasses is no longer the model but the sensor and power budget around it. Every slide is an attempt to buy always-on awareness at a power cost the frame can carry - 5mW for ambient vision, under 100mW for gaze, a 1.7mm³ eye module, cloud offload for anything heavy. Meta is fighting the same battle with a paywalled assistant and its own silicon roadmap, while Samsung leans on the Galaxy ecosystem. Alibaba's differentiator is that the agent can already reach ride-hailing, delivery, ticketing and payments inside its own super-app economy, which is exactly the plumbing Western platforms lack.
Two caveats. Every number here is Qwen's own, measured in Qwen's conditions, including the 90% reduction in false triggers and the sub-one-degree gaze accuracy, and none of it has been independently tested yet. And the timeline is split: the new hardware lands on a flagship agent-glasses model later this year, while the already-shipping Qwen S1 and G1 get the agent software layer through an OTA update, without the eye tracking, the 5mW vision chip or the 3DoF display.
We track every model, spec and release date in our smart glasses benchmark, and the Qwen line is in there alongside Meta, Samsung, XREAL, Rokid and the rest.
Source: Alibaba Group ↗
Interactive guide — prices, displays, cameras, battery.
Share this story





