About the role
About Humyn Labs Humyn Labs builds the intelligence layer for physical-world AI. We convert human action — sight, sound, movement, touch — into training-grade multimodal data for embodied AI, captured across 20+ countries in the real environments where physical AI actually deploys, not the labs where it is built. Our egocentric corpus ships to robotics and world-model teams as labeled data: 6-DoF head pose, metric 3D hand keypoints, depth, object tracks. The labels are the product. Their accuracy is the company. The opportunity Our auto-annotation stack currently runs on public off-the-shelf models — for head pose, hand keypoints, depth, and object detection and segmentation. Every one of them was trained on data that does not look like ours, and every one has a documented failure mode on our captures: wide-FOV fisheye rigs, gloved hands, heavy occlusion, motion blur, ego-motion, real workshops instead of clean scenes. We have measured those failures precisely. We are hiring you to remove them — by training our own models, on our own data, and putting them in production behind a gate. This is not a research role adjacent to the pipeline. You own the model layer of the pipeline. What you'll own • Metric 3D hand pose. Build the in-house hand model that is stable in the world frame, and decide the representation (MANO topology, 6D vs axis-angle) rather than inheriting it. • Egocentric pose that survives the camera. We have narrowed the suspects (rectification residual, FOV, shutter type) and currently route between several public SLAM / VIO systems per clip. Replace routing-by-heuristic with a model that makes any capture device deliverable — including learned despiking and drift correction. • Ground truth where we have none. Today we validate pose without ground truth (loop-closure drift, static-window jitter). Stand up real GT — fiducial-based (AprilTag / ArUco), gravity-aligned via IMU — under a hard constraint: our subjects are real workers doing their own jobs, so any protocol must add zero operator burden. • Our own depth model. Public depth models do not survive our rigs. We need a depth model of our own: metric, stereo-native, valid across the full frame on wide FOV, and cheap enough to run on every video hour we deliver. Own the training data, the architecture call, and the accuracy-versus-throughput trade. • Learned QC instead of thresholds. Our gates are hand-tuned numbers: epipolar residual, depth QC metrics, MCAP QA checks. Train error-detection models that flag a bad label before it reaches a customer, per clip, and quantify their catch rate against known-bad deliveries. • Training and serving, both. Train on AWS and ship on AWS: batch GPU orchestration, throughput tuning, daily delivery SLAs. Cost per labeled video hour is one of your numbers. What we're looking for • MS / PhD in computer vision, robotics or ML — or equivalent published research output. • You have trained a perception model that beat an off-the-shelf baseline on a real domain and shipped it. Data curation, loss design and eval decisions were yours. • Deep hands-on work in at least two of: 3D hand / human pose reconstruction, stereo or monocular depth, VIO / SLAM, open-vocabulary detection and segmentation, video VLMs. • Multi-view geometry is fluent, not looked up: rectification, epipolar residual, Q matrix and depth sign conventions, triangulation, world vs camera frame. • You can read a calibration report and say whether the problem is the camera or the model — and be right. • Strong Python and PyTorch; you own training and inference code, not notebooks. • Production inference experience on AWS: containerization, batch pipelines, orchestration. • 4+ years on problems in this space. Nice to have • Publications in egocentric vision, 3D hand pose, VLA / VLM models, or robot learning. • IMU-synced multimodal data, gravity alignment, camera–IMU extrinsics. • Robotics data formats: MCAP / Foxglove, ROS, LeRobot, RLDS. • Worked with Ego4D, EgoExo4D, Open X-Embodiment or comparable large egocentric corpora. • Open-source contributions to vision or robotics projects. First 60 days • Day 30. Reproduce our label QC numbers end to end and tell us which of the four label types is costing us the most, with evidence. • Day 60. One in-house model beating its public baseline on our eval set, on that label type. Compensation • No fixed budget — compensation is benchmarked to industry standard for the profile.