Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation
This paper presents a modular training-free framework for zero-shot, language-guided robotic manipulation in semi-structured environments. The architecture bridges the gap between high-level reasoning and low-level kinematics by decomposing the vision-action pipeline into three stages: visual perception, semantic interpretation, and task execution.

Evidence notes
- The modular training-free framework separates visual perception, semantic interpretation and task execution for language-guided manipulation.
- Experiments use the Robotiq Hand-E Adaptive Gripper in semi-structured zero-shot tasks.
- The design connects high-level semantic reasoning to conventional kinematic execution without end-to-end policy training.
Company context
Canadian collaborative-robot tooling company supplying adaptive grippers, vacuum grippers, force/torque sensing and related cobot components.