Skip to content
Thinking in Systems
ai-native #robots
Initializing search
infdahai/my-blog
Thinking in Systems
infdahai/my-blog
Welcome to MkDocs
Notes
Notes
DeepSeek Harness 与 Pi Agent 的架构分叉
Harness 到底怎么做
SGLang 通往 Maintainer 的路径
SGLang 架构 01 · 进程模型与请求环
SGLang 架构 02 · Scheduler 循环
My Knowledge Workflow
数据 infra 人员要求
ai-native #robots
Papers
Papers
EGO 数据线 · A · 分析蓝本(可执行)
EGO 数据线 · B · L2 表示/对齐深挖 —— ego 数据线 × 统一动作空间接口
EGO 数据线 · Danfei Xu(徐丹飞)组主线 roadmap
EGO 数据线 · 三大 EGO 数据集标注层次对照表
EGO 数据线 · EGO 数据集标注层次对照(扩展版,6数据集)
EGO 数据线 · 索引与梳理(消化卡)
EGO 跨本体桥 · 广度 roadmap + 精度 Insight
待精读队列
COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones
Zero-Shot Transfer of Force Map Estimation Across GelSight Mini Sensors
Performance-guided Task-specific Optimization for Multirotor Design
Enhancing Sim2Real Transfer for Torque-Controlled Robots through Real2Sim Dynamics Estimation and Reinforcement Learning
Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
Fast Generative Grasping via Lie Group-Constrained MeanFlow
TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks
Cross-Embodiment Robot Manipulation via a Unified Hand Action Space
CRAFT: Video Diffusion for Bimanual Robot Data Generation
Closing the Loop on the Poppy Humanoid: Bipedal Locomotion with Linear-Quadratic Control and Learned Cost Functions
Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation
Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
Anytime Global Tensor Motion Planning
PRISM: Projection-Integrated Sampling-Based MPC with Bayesian Cost Tuning for Bimanual Manipulation
The Embodiment Gap in Robot Foundation Models
VirTooS: A ROS 2 - Unity Virtualization Toolkit for Fleet Management of Autonomous Mobile Robots
Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies
UniDexTok: A Unified Dexterous Hand Tokenizer from Real Data
WAM-TTT:从 VLA、人类视频模仿到部署时技能记忆
Trust-Aware Sequential Decision Making and Rollout Planning for Resilient Multi-Robot Systems
Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts
Scalable Distributed Simulation-Based Testing for Automated Driving Systems
Beyond Shallow-Water Photorealism: Physically and Sensor-Grounded Simulation for Deep-Sea Robotics
RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience
Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation
Koala Gripper: Co-designing Robotic Grippers and Data-Capture Devices for Scaling Dexterous Manipulation Learning
SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
UNRIO: Uncertainty-Aware Velocity Learning for Radar-Inertial Odometry
Contact-Rich Robotic Manipulation in Construction via Zero-Shot Learning: A Diffusion Policy-Guided Adaptive Control
$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation
Physics Filtering Favors the Generalization of Robot Learning
GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Robots Need More than VLA and World Models
Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression
Cross-Platform Benchmark of Neural 3D Reconstruction for Autonomous Laboratory Robots
kimOpenVLAOpenSourceVisionLanguageAction2024a
Safety-Critical Bilateral Teleoperation for Omnidirectional Aerial Manipulation Using Force-Sensorless Haptic Feedback
In-Situ Reconstruction of the International Space Station Using 3D Gaussian Splatting and Astrobee
Extending Ground-Constraint LiDAR-IMU Calibration to Tilted Surfaces in a Continuous-Time Framework
The Missing Touch: Spatially Distributed Tactile Feedback Brings Teleoperation Closer to Human Dexterity
Marine Autonomous Vehicle Fleet Scheduling to Maximise Scientific Impact
Co-training with Ego-centric Video and Demonstration for Robot Navigation Task
$μ_0$: A Scalable 3D Interaction-Trace World Model
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery
Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions
DreamLedger: Execution-Settled Credit Files for World-Model Imagination in Robot Decision Loops
Handroid: Bridging Dexterous Hand and Humanoid
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Rethinking Demonstration Unlearning in Imitation Learning for Robotics
Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections
Turning Video Models into Generalist Robot Policies
HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models
G0.5: One Autoregressive Stream for Robot Reasoning and Action
LAC: Linear and Angular Compliance for Humanoid Whole-body Control
Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs
Scene2Demo: Self-Evolving Embodied Data Generation via Object-Action Graph
TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes
Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation
ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation
Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
Three-Way Open-Set Detection for Robust Autonomous Navigation
GeoWAM: Visual Geometry World Action Models for Autonomous Driving
Keypose Exploration: Efficient Automatic Trajectory Labelling and Cross-Embodiment Policy Transfer
Geometric Entropy: When Trajectory Diversity Helps and Hurts in Imitation Learning
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
Robust Slip Detection and Material Classification via Spatiotemporal Transformers on a Uniformly-Illuminated Visuo-Tactile Sensor
ROS2SmolVLA: Enabling Small Vision-Language-Action Models for Integration into Industrial-Grade Lightweight Robots
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Design of a Biomimetic Joint-Covering Skin with Tissue-Like Structure to Enhance Proprioception in a Musculoskeletal Humanoid
Longitudinal Robot Learning from Demonstration with Care Providers in a Home Environment
CSymPlan: Certified Symbolic Planning and Control for High-DOF Manipulators
A Statistical Audit of Physical AI Benchmark Redundancy
YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale
Reward-Free Continual Adaptation for Resilient Space Robots
UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models
CRESSim-Neo: A Batched GPU Simulation Engine for Surgical Robotics and Robot Learning
One-Shot Learning from Demonstration of Contact-Rich Robotic Manipulation by Identifying Physical Interactions
Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
Fiber Optic Sensing Glove for High Performance Dexterous Manipulation Capture
Reproducible Vision-Guided 6-DoF Robotic Manipulator with a Mixed Stepper-Driver Architecture and Browser-Native Control
Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA
DESCENT: Directed Edge Scene Encoding for Airport Surface Movement Prediction
4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation
RoboMME-Interference: Benchmarking Robot Memory Under Interference
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
A Taxonomy of Construction Task Activities for Robot Workers
Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF
GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
ROS2 Connect: A new ROS2 over WAN Solution
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
Diversity You Can Actually Measure: A Fast, Model-Free Diversity Metric for Robotics Datasets
HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation
Latent Chain-of-Thought World Modeling for End-to-End Driving
LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories
Video2DoorTraversal: Push Door Traversal via Simulated Door Twins
One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics
Model-Free Adaptive Parameter Tuning for Efficient Multi-Robot Warehouse Operations
Wave-Based Bilateral Teleoperation between Nonlinear Manipulators with Direct Contact Force Feedback
Human vs. Teleoperated Robots in Vineyard Management: A Simulation-Based Analysis of Travel Speed, Routing, and Task Performance
BehaviorWorldGen: Closing the Loop between Action Models and World Simulators via Controllable Behavior-Aware Structured World Generation
CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting
EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning
EgoNav: Bridging Learned Waypoints and Geometry-Aware Local Control for Robust Indoor Navigation
KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
From Foundation to Application: Improving VLA Models in Practice
$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models
Revisiting the "Push-T" Robot Manipulation Task with Agentic Robotics
SkyDrive: Learning to Drive in a New City from Aerial Traffic Monitoring
Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Data Analogies Enable Efficient Cross-Embodiment Transfer
Trajectory-Level Continuous Action Representation for Robotic Manipulation
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models
Gripper-aware Vision Language Action Models
GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM
AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning
EaDex: A Cross-Embodiment Dexterous Manipulation Framework from Low-Cost Demonstrations
InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation
SUPER ODOMETRY 2.0: Resilient Odometry via Hierarchical Adaptation
Attune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles
The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction
SurgSync: Time-Synchronized Multi-Modal Data Collection Framework and Dataset for Surgical Robotics
TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation
专题 · 多样性/覆盖 的量化与校准
专题延伸 · 统一动作空间 / 免重定向(以 UCAG-P 为锚)
企业数据集研究 · 首期产业勘查(具身头部公司)
具身智能 · 数据方向速览 2026-08-27
具身智能 · 数据集研究方向速览 2026-08-27
具身智能论文速览 2026-08-26 · 完整版
具身智能论文速览 2026-08-26 · 新增13篇
具身智能论文速览 2026-08-26
周报 · 数据×训练 方向(2026-08-27)
数据×训练 · 研究纲领(方向把握)
数据×训练 · 证据清单(按四层栈归位)
数据性质 × 有效训练 · 广度 roadmap + 精度 Insight
深度精读 · EGO/UMI 数据引擎批次(8篇)
深度精读 · 今日 · 跨本体 / 触觉 / 视频 / 合成数据 · 2026-08-27
深度精读 · 今日推送(数据×训练方向)· 2026-08-27
深度精读 · 数据方向剩余批次(17篇)
深度精读 · 数据方向 批次2(6篇)
深度精读 · 数据集方向两篇旗舰
Archive
Archive
2026
Categories
Categories
AI Engineering
ai-native #robots
Comment: Website: https://openvla.github.io/
Back to top