Track recent robotics research publications, benchmarks, datasets, model releases, and technical papers. Browse dated, source-backed robotics company events.
· research publication · ENGINEAI · PM01 · Humanoid
Evidence notes
The PM01 completed the full monkey-bar sequence in 14 of 15 hardware trials across three bar configurations and reached brachiation speeds up to 0.5 m/s.
A separately trained policy completed ten passes beneath randomly placed 2 cm wooden bars without contact, recovering to a stable stance after each crossing.
The research configuration replaced PM01 hands with passive hooks and mounted a RoboSense E1R lidar; after one goal position was supplied, the traversal sequence ran autonomously.
· research publication · Alibaba Group · Foundation Models
Evidence notes
ABot-Recon accepts monocular RGB video without requiring depth sensors or pre-calibrated camera parameters.
On KITTI-02 at 504 by 280 resolution, ABot-Recon reaches 24.45 frames per second with about 6.71 GB peak memory; the stated minimum GPU is a GTX 1080 Ti.
Inference code, evaluation scripts and pretrained weights were public at announcement; no future training-code release is treated as completed.
WALL-SS generated action-conditioned streaming rollouts up to 60 seconds while retaining recent history at fine resolution and compressing older state under a bounded memory budget.
Across 600 matched simulated and physical rollouts on six manipulation tasks, task-level success rates had mean absolute error of 0.062 and Pearson correlation of 0.93.
A co-trained action expert controlled a physical bimanual platform across seven tabletop tasks and reached average Task Progress of 69.1, ahead of the three reported comparison policies.
· research publication · Perceptron · Foundation Models
Evidence notes
Inputs can include images, video, language, robot state and previous actions; outputs include grounded visual responses, task progress and robot actions.
Perceptron reports training across more than 35 robot systems, 100,000 hours of robot experience, one million hours of general video and three trillion multimodal tokens.
Weights, training and inference code, the technical report and LeRobot support were released publicly.
· research publication · Lightwheel · Data Infrastructure
Evidence notes
The collection spans seven environment categories, 128 scene classes, more than 15,000 collection scenes and more than 15,000 tasks.
Head- and wrist-mounted views are paired with depth, hand-pose, full-body-pose and event-level semantic annotations for manipulation, coordination and task-sequence learning.
The first tranche is distributed through Hugging Face and AtomGit, and Lightwheel said the project will be donated to the OpenAtom Foundation for incubation.
· research publication · Generalist · GEN-1 · Foundation Models
Evidence notes
Across 10 tasks, Generalist reports 59% average success from one 3–12 second in-context demonstration and 83% after 10 gradient steps on five minutes of task data.
The multimodal model maintains 30 seconds of video context alongside sensor, language and proprioceptive inputs, and produces action trajectories at 100 Hz.
The release also shows prompts transferring from simulation to a physical robot and, in some cases, from human-hand demonstrations to immediate robot execution.
· research publication · Skild AI · S1 · Foundation Models
Evidence notes
S1 applies in-context learning to seen and unseen manipulation tasks, including plant potting, pancake preparation, pour-over coffee and kit assembly.
Skild reports unseen tasks lasting up to 10 minutes and a sevenfold improvement over language-only prompting on its evaluation set.
The official release demonstrates transfer from a human video example to robot execution, including an 11-minute demonstration-to-execution plant-potting workflow.
A new episode format with topic-group chunking reduced storage by about 68% and made sample reads approximately 2.9 times faster.
An Airflow-based ingestion architecture increased throughput from 14,000 to 440,000 episode-hours per week, cutting million-hour processing from roughly 16 months to under three weeks.
Warehouse-backed curation and memory-mapped tables reduced time to first training batch from about 48 hours to under one minute across a 43-million-episode dataset.