News & Releases

Humano OpticPro

Human-Centric Ciew for Robotics Augmentation

Executive Summary

Egocentric AI glasses are becoming a central platform for spatio-temporal intelligence and human-perspective transfer learning. By combining camera streams, eye tracking, microphones, inertial sensing, and on-device perception, these systems allow action models to observe the world from a human point of view. However, most wearable systems are still framed around recording, perception research, and offline scene understanding.

For industrial robotics, especially construction, a more active capability is required where such glasses do not only answer What did the user see? but also What is the user attending to, where is it in the world, what component does it represent, what are its dimensions, and how can this become a digital-twin and/or robot task target? Humano OpticPro addresses this gap through an attention-guided spatio-temporal pipeline.

The system combines gaze, IMU head orientation, voice commands, hand tracking, object localization, 3D mapping, SLAM, fiducial calibration, and VLM-based control to infer human attention and intent in real time. This allows the platform to support hands-free site surveying, architectural layout generation, visual-load estimation, ergonomic assessment, and robot-ready task grounding. In human-robot environments, the glasses convert the natural human behaviour of looking, speaking, and interacting into structured 3D information that guides inspection, navigation, manipulation, and shared autonomy.

System Specifications

Humano OpticPro is a closed-loop multimodal perceptual system that converts human intent into actionable 3D intelligence.

[1] The wearable intent layer captures what the user is doing through eye tracking, scene video, IMU-based head motion, microphone commands, and hand tracking. [2] The metric geometry layer provides stereo depth, dense depth maps, SLAM/VIO, and optional global positioning to understand where objects and surfaces exist in physical space. These streams are combined in the fusion and calibration core, where gaze is projected from the eye camera into the stereo camera, corrected using fiducial calibration, smoothed over time, and fused with depth to create a stable metric 3D target. That target is then passed to the foundation perception layer [3], where SAM 2 segmentation, object localization, and VLM-based understanding identify the surface, object, or scene component being referenced.

[4] The task output layer turns this perception into usable results, such as measurements, component capture, 2D layouts, project memory, robot-ready targets, and exports for BIM, USD, MJCF, dashboards, reports, among others. [5] The ergonomics and HRI layer runs alongside the pipeline, using gaze, posture, task sequence, and robot interaction data to estimate attention, visual load, shared-autonomy needs, and digital-twin updates. [6] The output layer results in a task-oriented active-perception and control bridge between the human, the physical environment, the digital twin, and the robot.

Highlights

  • 2D gaze, IMU head orientation, task focus, and spatial interactions for attention analysis and visual-load estimation.
  • Transforms what humans look at into 3D targets for inspection, navigation, manipulation, and shared autonomy.
  • Fiducial-based calibration/cataloguing and spatio-temporal photogrammetry for high-fidelity digital twinning.
  • Voice-controlled ergonomic actions for segmentation, dimensioning, and architectural layouts.
  • Hand tracking, object localization, 3D depth mapping, SLAM, and VLM-assisted control.