Robot-GSTGeometry-aware spatial-temporal robot
policy representation and evaluation

Simulation and evaluation before acting.
We reconstruct robotic scenes to reason about future states
and verify actions and outcomes before execution.

Represent the worldReason about future statesVerify actions and outcomes
Problem & motivation

What happens after the action?

Reliable manipulation requires geometric understanding, post-action reasoning and evaluation before execution.

01 / STATE REASONING

Beyond the action proposal

A plausible action does not ensure the intended final object state.

02 / EVALUATION

Before a real-world trial

Real-world trials are costly. We evaluate candidate behaviours in a reconstructed environment before rollout.

03 / GEOMETRY

Ground affordances in 3D

We ground language-guided keypoints in 3D geometry to support feasible manipulation.

Our idea: a geometry-aware simulator that verifies actions and outcomes before execution.

Method

A simulator for reasoning and evaluation

We integrate scene reconstruction, spatio-temporal reasoning and state-based trajectory planning in one robotic environment.

Robot-GST connects a reconstructed Gaussian environment, reasoning and prediction, and simulation with real-world manipulation.
Robot-GST overview · From RGB-D and language to simulation and evaluation before acting.
01
Represent

Gaussian-SAM3D simulation and evaluation environment

We combine 3D Gaussian Splatting and SAM3D with RGB-D observations to reconstruct scenes at real-world scale, aligned with the robot.

Appearance

Gaussian rendering provides photorealistic visual observations.

Geometry

Explicit scene and object geometry supports collision and contact reasoning.

Robot kinematics and physics/contact simulation govern the rollout; rendering supplies the observations.

RGB-D observations, segmentation and semantic features support Gaussian-SAM3D reconstruction.
Reconstruct the scene, then align its geometry with the robot.
02
Reason

Turn language into relational geometry

Task-relevant numbered keypoints on toy ducks and the basket.
Task-relevant keypoints define geometric constraints.

We lift task-relevant keypoints into metric 3D using depth, then encode proximity, alignment and relative pose as task constraints.

  1. 01

    3D keypoints + language task

    Identify task-relevant objects and relations.

  2. 02

    Relational constraints

    Specify stage goals and path requirements.

  3. 03

    Stage target pose

    Optimize the end-effector pose and gripper parameters.

03
Predict & verify

Plan with the object’s final state in mind

We predict final object states from candidate release poses, filter them with geometric checks, and plan trajectories toward the selected state.

Keypoint constraints and object-state reasoning lead to planned grasp, move and release stages in the reconstructed environment.
From task constraints to object-state prediction, motion planning and simulated rollout.
Candidate trajectory→Simulate→Evaluate

Executable motion

Inverse-kinematics feasibility and path requirements.

Collision safety

Geometric clearance along the candidate trajectory.

Task completion

The final object state satisfies the task goal.

If any check fails

Revise the plan and repeat simulation before execution.

If all checks pass

Execute the verified plan on the real robot.

Experiments

From reconstructed scenes to real manipulation

Our full task rollouts pair simulation with real-world manipulation.

Place a cube into a box

We grasp, transfer and place the cube inside the box without collision.

Real-world success rate80% → 90%Without → with replanning
Simulation
Real world

Full sequences, retimed for visual alignment.

From failure to revision

Completing the motion is not enough

Part of the toy remains outside the basket, failing the full-containment goal.

We revise the policy and evaluate it in simulation before execution, achieving successful packing.

Revise → Simulate and evaluate → Execute

incomplete toy packing followed by revised policy evaluation in simulation and successful real-world packing.
From the demonstrated packing failure to successful execution after replanning.
Quantitative results

Evidence for simulation-based evaluation

We evaluate reconstruction quality, sim–real consistency and manipulation success.

01

Simulation reflects real-world success trends

Our simulator captures real-world success trends across tasks, supporting evaluation before execution.

simulation and real-world success rates across tasks and over task progress, with higher replanning curves.
Left: correspondence between simulated and real-world success. Right: success over task progress; replanning improves success in both environments.
02

Evaluation and revision improve manipulation reliability

Our evaluation-and-replanning loop raises mean real-world success by 24.7 percentage points.

Without replanning55.3%
With replanning80.0%

Mean real-world success
Three-task average

Task success rates (%)
TaskWithout replanningWith replanning
SimulationReal worldSimulationReal world
Cube placing80.080.0100.090.0
Toy packing60.040.0100.080.0
Duck rearrangement60.046.080.070.0
Average66.755.393.380.0
03

High-quality observations from reconstructed scenes

We achieve the best SSIM and LPIPS, and second-best PSNR among the compared methods.

Ground truth, POGS, Embodied GS and Robot-GST rendering comparisons across robotic scenes.
Reconstruction quality across robotic scenes.
Rendering performance
MethodPSNR ↑SSIM ↑LPIPS ↓
GScream17.820.5900.560
VR-GS24.130.8330.320
Decoupled GS27.320.9050.300
POGS19.660.8690.098
D3DGS12.050.3990.230
Embodied GS14.950.3890.360
Robot-GST (ours)24.640.9630.044

PSNR in dB; LPIPS uses VGG. Bold: best. Underline: second-best.

Research summary

Geometry-aware representation,
state-aware execution

We build a Gaussian-SAM3D robotic environment that turns visual observations and language instructions into geometric constraints, predicted object states and executable motion.

By simulating candidate behaviours before execution, we evaluate both trajectory feasibility and task outcomes. Across three manipulation tasks, our framework connects geometric reasoning with more reliable real-world execution.

Figure detail