Skip to content

Papers

具身智能(Embodied AI)方向的论文阅读笔记。每篇论文一个 Markdown 文件,由 Obsidian Citations 插件根据 90_Templates/paper-note.md 自动生成骨架,文件名 使用 Better BibTeX 生成的 citekey。

PDF 原件由 Zotero 管理,存储在 iCloud (~/Library/Mobile Documents/com~apple~CloudDocs/Zotero_Data/), 不进入 Git,不在博客发布。

索引

  • [[fengWAMTTTSteeringWorldAction2026|WAM-TTT:从 VLA、人类视频模仿到部署时技能记忆]]
  • [[kimOpenVLAOpenSourceVisionLanguageAction2024a|OpenVLA: An Open-Source Vision-Language-Action Model]]
  • [[chenDoWorldModelsReallyFollowActionsDiagnosingAligning2026|Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning]]
  • [[overbeekOneShotLearningDemonstrationContactRichManipulationIdentifying2026|One-Shot Learning from Demonstration of Contact-Rich Robotic Manipulation by Identifying Physical Interactions]]
  • [[zhangGaussianWAMDistillingGeometrySemantics3DGaussianFieldsWorld2026|GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models]]
  • [[zhangGripperAwareVisionLanguageActionModels2026|Gripper-aware Vision Language Action Models]]
  • [[peifferFiberOpticSensingGloveHighPerformanceDexterousManipulation2026|Fiber Optic Sensing Glove for High Performance Dexterous Manipulation Capture]]
  • [[tejeroNVIDIACosmosHDreamsRealTimeGenerativePhysics2026|NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics]]
  • [[maRobustSlipDetectionMaterialClassificationSpatiotemporalTransformersUniformly2026|Robust Slip Detection and Material Classification via Spatiotemporal Transformers on a Uniformly-Illuminated Visuo-Tactile Sensor]]
  • [[choiPonderPouncePretrainedMLLMEpisodeContextEngineControl2026|PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control]]
  • [[yangTrajectoryLevelContinuousActionRepresentationManipulation2026|Trajectory-Level Continuous Action Representation for Robotic Manipulation]]
  • [[caoTrActBridgingControlVisualPredictionVisualTracks2026|TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks]]
  • [[haoHierarchicalSkillRetrievalDataEfficientAdaptationVisionLanguage2026|Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models]]
  • [[liDreamLedgerExecutionSettledCreditFilesWorldModelImagination2026|DreamLedger: Execution-Settled Credit Files for World-Model Imagination in Robot Decision Loops]]
  • [[luGeoWAMVisualGeometryWorldActionModelsAutonomousDriving2026|GeoWAM: Visual Geometry World Action Models for Autonomous Driving]]
  • [[leeActIntentDistillingBehaviorIntentVisionLanguageAction2026|Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models]]
  • [[orsulaRewardFreeContinualAdaptationResilientSpace2026|Reward-Free Continual Adaptation for Resilient Space Robots]]
  • [[mandischerROS2SmolVLAEnablingSmallVisionLanguageActionModelsIntegration2026|ROS2SmolVLA: Enabling Small Vision-Language-Action Models for Integration into Industrial-Grade Lightweight Robots]]
  • [[mikiDesignBiomimeticJointCoveringSkinTissueLikeStructure2026|Design of a Biomimetic Joint-Covering Skin with Tissue-Like Structure to Enhance Proprioception in a Musculoskeletal Humanoid]]
  • [[zhouThinkOnlyWhenNeededPromptAuthorityControlSelective2026|Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation]]
  • [[chenPointingVLATypedSpatialGroundingInterfacesVisionLanguage2026|Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation]]
  • [[zhaoInstructMoveTextIndispensableBenchmarkInstructionFollowingManipulation2026|InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation]]
  • [[narendraCSymPlanCertifiedSymbolicPlanningControlHighDOFManipulators2026|CSymPlan: Certified Symbolic Planning and Control for High-DOF Manipulators]]
  • [[osterbergUniMemUnifyingMultimodalMemoryControlVisionLanguageAction2026|UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models]]
  • [[liuTriplet2TrackHierarchicalSystemObjectCentricRepresentationsReliableLong2026|Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation]]
  • [[pereraReproducibleVisionGuided6DoFManipulatorMixedStepper2026|Reproducible Vision-Guided 6-DoF Robotic Manipulator with a Mixed Stepper-Driver Architecture and Browser-Native Control]]
  • [[liuPhysicalAgenticAIArchitectureOrchestratingCrewLLMs2026|Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs]]
  • [[zhouZeroWAMContextWorldActionModelingHumanVideos2026|Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization]]
  • [[bukhariFastGenerativeGraspingLieGroupConstrainedMeanFlow2026|Fast Generative Grasping via Lie Group-Constrained MeanFlow]]
  • [[teamOnePolicyManyEmbodimentsUnifiedCameraCentricAction2026|One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation]]
  • [[wuR3TrainingReasonNaturalLanguageReinforcementLearning2026|\(R^3\): Training Robots to Reason in Natural Language via Reinforcement Learning]]
  • [[zhangMAVLAMultiArmVisionLanguageActionModel2026|MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization]]
  • [[zhouTacForcingStreamingActionGenerationExecutionTimeTactileFeedback2026|TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback]]
  • [[louLMXExplainableActionModelingProgressEventUncertainty2026|LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation]]
  • [[danPRISMProjectionIntegratedSamplingBasedMPCBayesianCost2026|PRISM: Projection-Integrated Sampling-Based MPC with Bayesian Cost Tuning for Bimanual Manipulation]]
  • [[jiangGaussianDreamEfficient3DGaussianWorldModelingManipulation2026|GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation]]
  • [[wangEgoNavBridgingLearnedWaypointsGeometryAwareLocalControl2026|EgoNav: Bridging Learned Waypoints and Geometry-Aware Local Control for Robust Indoor Navigation]]
  • [[jangRAVLARetrievalAugmentedVLATestTimeAdaptation2026|RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation]]
  • [[liuConfALWMConfidenceGuidedActiveLearningActionConditioned2026|ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models]]
  • [[liuLACLinearAngularComplianceHumanoidWholeBodyControl2026|LAC: Linear and Angular Compliance for Humanoid Whole-body Control]]
  • [[sakibTaxonomyConstructionTaskActivitiesWorkers2026|A Taxonomy of Construction Task Activities for Robot Workers]]
  • [[sarowarGaussVLAGeometryAwareSpatialReasoningVisionLanguageAction2026|GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model]]
  • [[schottROS2ConnectNewROS2WANSolution2026|ROS2 Connect: A new ROS2 over WAN Solution]]
  • [[korotkineExtendingGroundConstraintLiDARIMUCalibrationTiltedSurfaces2026|Extending Ground-Constraint LiDAR-IMU Calibration to Tilted Surfaces in a Continuous-Time Framework]]
  • [[xiongSkyDriveLearningDriveNewCityAerialTrafficMonitoring2026|SkyDrive: Learning to Drive in a New City from Aerial Traffic Monitoring]]
  • [[ouCRESSimNeoBatchedGPUSimulationEngineSurgicalRobotics2026|CRESSim-Neo: A Batched GPU Simulation Engine for Surgical Robotics and Robot Learning]]
  • [[moormanLongitudinalLearningDemonstrationCareProvidersHomeEnvironment2026|Longitudinal Robot Learning from Demonstration with Care Providers in a Home Environment]]
  • [[lesaniSaliencyDepthConditioningZeroShotSegmentationCommunicationTower2026|Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery]]
  • [[francosTrustAwareSequentialDecisionMakingRolloutPlanningResilient2026|Trust-Aware Sequential Decision Making and Rollout Planning for Resilient Multi-Robot Systems]]
  • [[liChooseYourGameWiselyMeasuringGameTheoreticStructures2026|Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions]]
  • [[udekweHumanVsTeleoperatedVineyardManagementSimulationBasedAnalysis2026|Human vs. Teleoperated Robots in Vineyard Management: A Simulation-Based Analysis of Travel Speed, Routing, and Task Performance]]
  • [[liuScene2DemoSelfEvolvingEmbodiedDataGenerationObjectAction2026|Scene2Demo: Self-Evolving Embodied Data Generation via Object-Action Graph]]
  • [[huangUNRIOUncertaintyAwareVelocityLearningRadarInertialOdometry2026|UNRIO: Uncertainty-Aware Velocity Learning for Radar-Inertial Odometry]]
  • [[zhangScaRFSLAMScaleConsistentReconstructionFeedForwardModels2026|ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM]]
  • [[jajooRegularizedLatentDynamicsPredictionIsStrongBaselineBehavioral2026|Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models]]
  • [[linPoint2PoseOcclusionRecovering6DPoseTracking3DReconstruction2026|Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers]]
  • [[navasardyanStatisticalAuditPhysicalAIBenchmarkRedundancy2026|A Statistical Audit of Physical AI Benchmark Redundancy]]
  • [[bargelliniEnhancingSim2RealTransferTorqueControlledThroughReal2SimDynamics2026|Enhancing Sim2Real Transfer for Torque-Controlled Robots through Real2Sim Dynamics Estimation and Reinforcement Learning]]
  • [[boraMacroOperatorGenerationPredicateSelectionTAMPOperatorLearning2026|Macro-Operator Generation and Predicate Selection for TAMP Operator Learning]]
  • [[zhouImitatorGameBenchmarkingImitativeAbilityBeyondActionPrediction2026|The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction]]
  • [[wangBehaviorWorldGenClosingLoopBetweenActionModelsWorldSimulators2026|BehaviorWorldGen: Closing the Loop between Action Models and World Simulators via Controllable Behavior-Aware Structured World Generation]]
  • [[ibrahimovContactRichManipulationConstructionZeroShotLearningDiffusion2026|Contact-Rich Robotic Manipulation in Construction via Zero-Shot Learning: A Diffusion Policy-Guided Adaptive Control]]
  • [[zhangGuardianBenchSameSceneInstructionContrastiveBenchmarkLatentContextual2026|GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI]]
  • [[kimSafetyCriticalBilateralTeleoperationOmnidirectionalAerialManipulationForce2026|Safety-Critical Bilateral Teleoperation for Omnidirectional Aerial Manipulation Using Force-Sensorless Haptic Feedback]]
  • [[kimSituReconstructionInternationalSpaceStation3DGaussianSplatting2026|In-Situ Reconstruction of the International Space Station Using 3D Gaussian Splatting and Astrobee]]
  • [[liuViTacPhysPhysicalPropertyAwareGraspingHumanVisualTactile2026|ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations]]
  • [[gellerScalableDistributedSimulationBasedTestingAutomatedDrivingSystems2026|Scalable Distributed Simulation-Based Testing for Automated Driving Systems]]
  • [[hajjahmadKoalaGripperCoDesigningGrippersDataCaptureDevices2026|Koala Gripper: Co-designing Robotic Grippers and Data-Capture Devices for Scaling Dexterous Manipulation Learning]]
  • [[tangVideo2DoorTraversalPushDoorTraversalSimulatedDoorTwins2026|Video2DoorTraversal: Push Door Traversal via Simulated Door Twins]]
  • [[maDECOWAMDecoupledWholeBodyWorldActionModelLegged2026|DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation]]
  • [[tranWaveBasedBilateralTeleoperationBetweenNonlinearManipulatorsDirect2026|Wave-Based Bilateral Teleoperation between Nonlinear Manipulators with Direct Contact Force Feedback]]
  • [[merandCoToGraspContactTopologyConditionedDexterousGraspSynthesisCanonical2026|CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning]]
  • [[zhuSCAPEScenarioConditionedSimulationAugmentedPolicyEvaluation2026|SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation]]
  • [[kotaMissingTouchSpatiallyDistributedTactileFeedbackBringsTeleoperation2026|The Missing Touch: Spatially Distributed Tactile Feedback Brings Teleoperation Closer to Human Dexterity]]
  • [[jingSoftVTBenchDeformationAwareVisuoTactileDatasetBenchmarkDeformable2026|SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation]]
  • [[endoOrienteeringProblemUncertainTimeVaryingRewardsFrameworkBenchmark2026|Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics]]
  • [[tangLabDexHierarchicalBenchmarkDexterousManipulationLaboratories2026|LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories]]
  • [[amorosZeroShotTransferForceMapEstimationGelSightMini2026|Zero-Shot Transfer of Force Map Estimation Across GelSight Mini Sensors]]
  • [[xieRevisitingPushTManipulationTaskAgenticRobotics2026|Revisiting the "Push-T" Robot Manipulation Task with Agentic Robotics]]
  • [[songHABITHumanAwareBehaviorInteractionTrainingDatasetManipulation2026|HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation]]
  • [[liOpenAoEOpenEgocentricManipulationDatasetToolchainEmbodied2026|Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning]]
  • [[zhouSurgSyncTimeSynchronizedMultiModalDataCollectionFramework2026|SurgSync: Time-Synchronized Multi-Modal Data Collection Framework and Dataset for Surgical Robotics]]
  • [[chenCRAFTVideoDiffusionBimanualDataGeneration2026|CRAFT: Video Diffusion for Bimanual Robot Data Generation]]
  • [[zhangDexoraOpenSourceVLAHighDoFBimanualDexterity2026|Dexora: Open-source VLA for High-DoF Bimanual Dexterity]]
  • [[guoWorldsOneDemoSyntheticDataEngineLearningOpen2026|Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation]]
  • [[wangLearningWhileDeployingFleetScaleReinforcementLearningGeneralist2026|Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies]]
  • [[agarwalCOBALTCrowdsourcingLearningCloudBasedTeleoperationSmartphones2026|COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones]]
  • [[marpallyACMEMultiCulturalMultiEmbodimentSocialNavigationDataset2026|ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset]]
  • [[wuSIEVEStructureAwareDataSelectionImitationLearningVLA2026|SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models]]
  • [[yaoDexVerseModularBenchmarkMultiTaskMultiEmbodimentDexterous2026|DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation]]
  • [[zhaoEaDexCrossEmbodimentDexterousManipulationFrameworkLowCost2026|EaDex: A Cross-Embodiment Dexterous Manipulation Framework from Low-Cost Demonstrations]]
  • [[fangUniDexTokUnifiedDexterousHandTokenizerRealData2026|UniDexTok: A Unified Dexterous Hand Tokenizer from Real Data]]
  • [[limHRDexDBPairedHumanDatasetCrossEmbodimentDexterousGrasping2026|HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping]]
  • [[wangEgoInfinityWebScale4DHandObjectInteractionData2026|EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning]]
  • [[zhaoAXISGrowableCommunityDrivenDataEngineScalableManipulation2026|AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation]]
  • [[lee0Scalable3DInteractionTraceWorldModel2026|\(μ_0\): A Scalable 3D Interaction-Trace World Model]]
  • [[ohkawaYUBIYieldingUniversalBidigitalInterfaceBimanualDexterousManipulation2026|YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale]]
  • [[sirigiriDiversityYouCanActuallyMeasureFastModelFree2026|Diversity You Can Actually Measure: A Fast, Model-Free Diversity Metric for Robotics Datasets]]
  • [[intelligence07SteerableGeneralistFoundationModelEmergentCapabilities2026|\(π_{0.7}\): a Steerable Generalist Robotic Foundation Model with Emergent Capabilities]]
  • [[wangVisionLanguageActionRoboticsSurveyDatasetsBenchmarksData2026|Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines]]
  • [[liTurningVideoModelsGeneralistPolicies2026|Turning Video Models into Generalist Robot Policies]]
  • [[wangKITEDecouplingKinematicsInteractionZeroShotCrossEmbodiment2026|KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation]]
  • [[luoPassiveVideoEditableExperiencePhysicallyGroundedExperienceSynthesis2026|From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence]]
  • [[yangDataAnalogiesEnableEfficientCrossEmbodimentTransfer2026|Data Analogies Enable Efficient Cross-Embodiment Transfer]]
  • [[karciniNeedMoreThanVLAWorldModels2026|Robots Need More than VLA and World Models]]
  • [[kunoCoTrainingEgoCentricVideoDemonstrationNavigationTask2026|Co-training with Ego-centric Video and Demonstration for Robot Navigation Task]]
  • [[liHandroidBridgingDexterousHandHumanoid2026|Handroid: Bridging Dexterous Hand and Humanoid]]
  • [[pisenoCloakZeroShotCrossEmbodimentManipulationMaskingEnd2026|Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA]]
  • [[xuMotionFocusedLatentActionEnablesCrossEmbodimentVLA2026|Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos]]
  • [[casasCrossEmbodimentManipulationUnifiedHandActionSpace2026|Cross-Embodiment Robot Manipulation via a Unified Hand Action Space]]
  • [[yuanEmbodiedR15EvolvingPhysicalIntelligenceEmbodiedFoundation2026|Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models]]
  • [[domaeEmbodimentGapFoundationModels2026|The Embodiment Gap in Robot Foundation Models]]
  • [[weiAmbientDiffusionPolicyImitationLearningSuboptimalDataRobotics2026|Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics]]
  • [[wuFoundationApplicationImprovingVLAModelsPractice2026|From Foundation to Application: Improving VLA Models in Practice]]
  • [[luKeyposeExplorationEfficientAutomaticTrajectoryLabellingCrossEmbodiment2026|Keypose Exploration: Efficient Automatic Trajectory Labelling and Cross-Embodiment Policy Transfer]]
  • [[holkAuditingInstructionTrajectoryMismatchesMultimodalDemonstrations2026|Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations]]
  • [[luoGeometricEntropyWhenTrajectoryDiversityHelpsHurtsImitation2026|Geometric Entropy: When Trajectory Diversity Helps and Hurts in Imitation Learning]]
  • [[jiaPhysicsFilteringFavorsGeneralizationLearning2026|Physics Filtering Favors the Generalization of Robot Learning]]
  • [[liSetSupervisedDiffusionPolicyLearningActionChunkingDiffusion2026|Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections]]
  • [[drudiVirTooSROS2UnityVirtualizationToolkitFleetManagement2026|VirTooS: A ROS 2 - Unity Virtualization Toolkit for Fleet Management of Autonomous Mobile Robots]]
  • [[zhaoSUPERODOMETRY20ResilientOdometryHierarchicalAdaptation2026|SUPER ODOMETRY 2.0: Resilient Odometry via Hierarchical Adaptation]]
  • [[chengRobustBimanualVisionLanguageActionModelsEmbarrassinglySimple2026|Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking]]
  • [[shenLD4WAMLearningLatentDynamicsHumanVideosWorldAction2026|LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models]]
  • [[garcesModelBasedReinforcementLearningHeterogeneousMultiTaskAssignment2026|Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts]]
  • [[tokekarModelFreeAdaptiveParameterTuningEfficientMultiWarehouse2026|Model-Free Adaptive Parameter Tuning for Efficient Multi-Robot Warehouse Operations]]
  • [[hanSRLMPCShapeAwareReinforcementLearnedModelPredictive2026|SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control]]
  • [[liRethinkingDemonstrationUnlearningImitationLearningRobotics2026|Rethinking Demonstration Unlearning in Imitation Learning for Robotics]]
  • [[guoRoboEditTurningHumanManipulationVideosScalableExperience2026|RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience]]
  • [[yuPRISMPrecisionContactRichRealWorldIndustrialSkill2026|PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing]]
  • [[zhaoNebulaVLADualFrequencyVisionLanguageActionModelGuide2026|NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation]]
  • [[longScalingManualGroundedApplianceManipulationDataSynthesisUnified2026|Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning]]
  • [[chenReflexEnablingFastPredictiveVisionLanguageActionModels2026|Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation]]
  • [[zhaoAdvDexLearningDexterousManipulationHumanDemonstrationsJointAligned2026|AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning]]
  • [[rongH2RBenchBenchmarkingHumanManipulationVideoGenerationWorld2026|H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models]]
  • [[zhouAttuneSelfAnnotationToolUnderstandingOperatorAttentionProfiles2026|Attune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles]]
  • [[liuG05OneAutoregressiveStreamReasoningAction2026|G0.5: One Autoregressive Stream for Robot Reasoning and Action]]
  • [[huangRynnValueScalingValueFoundationModelsTemporalDistance2026|RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance]]
  • [[linJEPAWAMLearningVisionLanguageActionPoliciesJoint2026|JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling]]
  • [[paulEgoOSCAREgocentricOpenSourceStereoCAptuReSystem2026|Ego-OSCAR: Egocentric Open source Stereo CAptuRe System]]
  • [[kimAreVisualPlaceRecognitionModelsRecognizingPlacesConditions2026|Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression]]
  • [[sanjayScalableLongHorizonPlanningStaggeredUpdatesLifelongMAPF2026|Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF]]
  • [[wangCrossTracerCrossEmbodimentNavigationVLAModelReasoningTrace2026|CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting]]
  • [[coumarAnytimeGlobalTensorMotionPlanning2026|Anytime Global Tensor Motion Planning]]
  • [[prutschDESCENTDirectedEdgeSceneEncodingAirportSurfaceMovement2026|DESCENT: Directed Edge Scene Encoding for Airport Surface Movement Prediction]]
  • [[xuVBVRProScalableVerifiableSuiteNativeVisualReasoning2026|VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning]]
  • [[arzaPerformanceGuidedTaskSpecificOptimizationMultirotorDesign2026|Performance-guided Task-specific Optimization for Multirotor Design]]
  • [[weiLearningAccelerateVisionLanguageActionModelsThroughAdaptive2026|Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching]]
  • [[rathiRoboMMEInterferenceBenchmarkingMemoryInterference2026|RoboMME-Interference: Benchmarking Robot Memory Under Interference]]
  • [[loukovitisModelAgnosticOpenSetAirAirVisualObject2026|Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception]]
  • [[loukovitisThreeWayOpenSetDetectionRobustAutonomousNavigation2026|Three-Way Open-Set Detection for Robust Autonomous Navigation]]
  • [[tanLatentChainThoughtWorldModelingEndEndDriving2026|Latent Chain-of-Thought World Modeling for End-to-End Driving]]
  • [[liuCLAPCrossEmbodimentVideoWorldModelsAreZero2026|CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators]]
  • [[krariMarineAutonomousVehicleFleetSchedulingMaximiseScientificImpact2026|Marine Autonomous Vehicle Fleet Scheduling to Maximise Scientific Impact]]
  • [[fangSpatialCrafterSingleImageWorldModelingGenerative3DProxies2026|SpatialCrafter: Single Image World Modeling with Generative 3D Proxies]]
  • [[qi4DSynthControllableProceduralWorldSynthesisDynamicEmbodiedSimulation2026|4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation]]
  • [[grimaldiBeyondShallowWaterPhotorealismPhysicallySensorGroundedSimulation2026|Beyond Shallow-Water Photorealism: Physically and Sensor-Grounded Simulation for Deep-Sea Robotics]]
  • [[liuTrapVLATrappingVisionLanguageActionModelsConfiguredFailure2026|TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes]]
  • [[chenClosingLoopPoppyHumanoidBipedalLocomotionLinearQuadratic2026|Closing the Loop on the Poppy Humanoid: Bipedal Locomotion with Linear-Quadratic Control and Learned Cost Functions]]
  • [[kimCrossPlatformBenchmarkNeural3DReconstructionAutonomousLaboratory2026|Cross-Platform Benchmark of Neural 3D Reconstruction for Autonomous Laboratory Robots]]