I am currently a graduate research assistant at the CRAFT Lab, Northeastern University, advised by Prof. Gilbert Yang Ye. My research focuses on safe, physically-grounded robot manipulation — building learned policies that respect safety constraints and genuinely exploit the sensory inputs they are given.
Before joining CRAFT Lab, I worked as a research assistant at SJTU IWIN-FINS Lab advised by Prof. Jianping He, where I co-designed the full-stack Fines robotic platform, covering hardware architecture, embedded control framework, and perception pipeline.
On the industry side, I built production robotics and automation systems. At Dinnar Automation, I designed automated optical inspection (AOI) equipment for Texas Instruments semiconductor products. At CloudMinds Robotics, I co-designed standardized multi-robot assembly workstations for harmonic-reducer production, with EtherCAT real-time control and digital-twin systems. At NIO Automobile, I developed sim-to-real transfer pipelines for autonomous driving.
A learned world model provides a powerful physical intuition for evaluating future states. But its effectiveness in continuous control also depends critically on how candidate actions are generated for model-based planning. Rather than solely asking how accurately a model can simulate the future, we ask: which candidate actions are worth evaluating in the first place? Existing planners typically search arbitrarily or use expert demonstrations only to initialize a sampling mean, discarding the expert's state-conditioned confidence. Properly guiding this search requires a robust action prior, yet current approaches often rely on independent visual encoders or large-scale VLMs to obtain one. We argue that this architectural bloat is unnecessary: the exact same data — and the learned representations of the world model itself — inherently encode the agent's action intuition. We introduce PRISM, a task-agnostic framework that extracts both from a single dataset while maintaining strict architectural simplicity. Building on a standard JEPA-style latent world model, PRISM attaches a lightweight MLP directly to its frozen encoder to predict a state-conditioned Gaussian prior. At plan time, PRISM fuses this prior into the planner's sampling distribution via a precision-weighted Product-of-Gaussians update. This parameter-free, closed-form integration steers the sampling process, making the prior confident where it is and ceding control where it is not. PRISM improves success rates by 35 percentage points over vanilla world-model-based MPC on Cube and 32 percentage points on PushT, without introducing significant inference overhead.
Contact-rich robotic manipulation has long relied on vision-based policies, but at the contact phase vision is often insufficient: the end-effector occludes the workpiece, and under commodity sensing the last millimeter is hard to close from vision alone. Force/torque sensing complements vision at exactly these moments, yet it is often unavailable at deployment: training datasets come from sensor-equipped platforms while target robots mount no force/torque sensor, and wrench estimates drift out of calibration. Existing force-conditioned policies treat force as an always-available input and do not address a missing sensor at test time. A natural remedy is to add an encoded wrench into the latent state of a vision world model, so that the model uses force when present and falls back to vision otherwise. In practice this backfires: on our force-augmented PushT benchmark (PushT-FT), planning success collapses from 41% to 15% once force is missing at test time. The collapse persists across fusion operators and even when missing force dominates the training data. What matters is the amplitude at which force is written into the shared latent. To address this, we propose ForceLeWM, a JEPA world model that writes force into the latent through a conservative per-dimension gate at a small learned amplitude, trained with a warmup alignment loss for the from-scratch setting. ForceLeWM restores planning to 35%, recovering about 75% of the lost planning success across 3 training seeds. A train-consistent normalization audit shows the perception benefit of force is task-specific: clear on FORGE peg insertion but not across other task families. We further deploy one frozen simulation-trained checkpoint end-to-end on two real Franka assembly tasks, gear meshing and peg insertion, with force supplied by a proprioceptive wrench observer estimated from joint torques rather than a dedicated force/torque sensor. When test-time force availability is uncertain, fusion should be conservative by design: admit force at a small learned amplitude when present, and degrade gracefully rather than collapse when absent.
In latent world models of the joint-embedding family, the planner learns about actions only through the model, and the model compresses each action block through a linear bottleneck. The compression is uneven over action directions, and along some directions it is total. We call the non-uniformity action-compression anisotropy. Existing treatments of action insensitivity modify the training objective and leave released, frozen models undiagnosed. This paper introduces an external protocol that diagnoses them. It reads per-direction sensitivity from end-to-end Jacobians, compares it with the response the environment exhibits along the same directions, and measures how much control the planner exerts along each. On four released checkpoints the protocol separates two kinds of compression. Compression learned by training is faithful, with sensitivity matching environment response at correlation 0.87 to 0.97. Compression imposed by the architecture is a blind spot, and one checkpoint is exactly insensitive to fifteen directions carrying about 8% of the environment's action response. The diagnosis supports a training-free intervention. Projecting executed actions onto the model's visible subspace raises cube manipulation success from 68.5% to 75.0%. The gain has a stated validity boundary. The projection helps only where the planner has no measured control over the deleted directions, and it costs up to 18.0 points where weak control survives. The state channel offers no such option. Projecting the planning cost onto the reachable subspace loses 18.7 to 25.3 points, and no tested rank recovers its baseline. A restriction is safe when the discarded component carries no planning signal, and that can be measured with everything frozen.
Vision-Language-Action (VLA) models have advanced robotic navigation by unifying perception and planning, yet most systems operate in an open-loop manner, generating plans without awareness of locomotion feasibility or the ability to adapt when execution becomes unsafe. This disconnect leads to cascading failures on challenging terrains and surfaces where pre-planned trajectories become infeasible. We propose Thinking on Its Feet (TOIF), a three-tier hierarchical control framework that closes the loop between low-level proprioceptive sensing and high-level semantic path planning for terrain-aware humanoid robot navigation. At the low level, an RL locomotion policy is augmented with a multi-detector anomaly system that evaluates terrain-specific feasibility and safety constraints using compact proprioceptive features. Detected anomalies are accumulated through leaky integrators and translated into structured natural-language feedback, enabling the high-level planner to replan over a semantic scene graph. At the mid level, a VLA model generates velocity commands that bridge high-level waypoints and low-level execution. We further introduce a risk-aware velocity modulation mechanism that proactively reduces speed based on real-time anomaly confidence. We validate TOIF in Isaac Lab simulation and on a physical humanoid robot across a range of challenging terrains. Results demonstrate that TOIF substantially improves navigation success rate and safety over open-loop baselines, and that the proprioceptive anomaly detectors transfer effectively to real-world deployment.
Unfolding a crumpled garment from an overhead view is ambiguous: a flat cloth and a stack of hidden folds can share the same silhouette. This paper studies a transparent-table dual-arm setup in which the robot observes both a top segmentation mask and a bottom contact mask. The main finding is that the bottom view is most useful when converted into fold-contact affordances rather than used as a generic second image. We derive fold residual, contact overlap, fold-contact frontier, and silhouette-edge maps from the dual masks, use them in an attention encoder, and further convert the frontier into a residual action prior. To reduce hand tuning, we parameterize the prior's cell scorer as an online linear, contextual-bandit-style update trained from reward, self-supervised unfolding progress, or their mixture. Across the completed logs, raw dual-view pixels improve long-horizon flatness over a top-only policy, affordance channels reduce folded area by 82.2% relative to raw dual masks in a fast screen, and a reward-bandit bottom-view scorer gives the best 21-run residual-prior ablation on within-panel geometry score. The results support a precise conclusion: bottom view helps cloth unfolding by exposing where upper folded layers meet grounded support, and reward-trained scoring is the most effective tested way to turn that information into manipulation targets.
Diffusion policies have become the dominant paradigm for learned robotic manipulation, yet they offer no intrinsic mechanism to enforce workspace safety constraints. TACS is a training-free method that injects barrier-function guidance into the DDIM denoising loop of action-space diffusion policies, applying repulsive guidance on the Tweedie-predicted end-effector positions; subsequent denoising steps coherently adapt rotation and gripper dimensions, achieving effective obstacle avoidance at 10× smaller guidance scales than post-hoc potential-field projection. On three robosuite manipulation tasks, TACS Pareto-dominates post-hoc filtering, with a statistically significant win-win on Lift (violations 6.6% → 0.32%, 20× reduction; p=0.015 across 5 seeds × 250 episodes) and zero measurable inference overhead versus the unmodified diffusion policy.
Odometry estimation remains a critical challenge for wheeled robots, as reducing its drift directly mitigates dependency on external localization systems. This paper proposes a distributed odometry framework for steerable wheels, named ICF-DO, which is applicable to both Steerable Wheeled Mobile Robots (SWMRs) and cooperative multi-single-wheel robot systems. The proposed method features low computational complexity and reduced drift, while demonstrating strong robustness in communication-restricted scenarios. Additionally, singularity can be processed in a distributed manner in the proposed framework. Experimental validation on a real physical SWMR platform demonstrates the effectiveness and practicality of the proposed method.
The vital component of the high-pressure pump is the high-pressure cylinder, which undergoes pulsating cyclic loads during operation. This exposure can lead to fatigue cracking and limit the pump’s overall performance. To address this issue, a swage autofrettage treatment method for the high-pressure cylinder is proposed based on the principles of autofrettage technology. The structure parameters of the mandrel are designed to optimize its role in the treatment process. Subsequently, a swage autofrettage simulation model is established, considering the Bauschinger effect to analyze the distribution of residual stress about the mandrel interference percentage. The results demonstrate that the swage autofrettage treatment improves the stress distribution in the high-pressure cylinder, causing the inner wall surface to undergo a circumferential residual compressive stress. Within a specific range, the Bauschinger effect becomes more significant as the mandrel interference percentage increases. This study provides insights for enhancing the durability and performance of high-pressure cylinders while adhering to mechanical engineering standards.
@article{sun2023simulation,
title={Simulation Research on Residual Stress of Swage
Autofrettage-processed High-Pressure Cylinder},
author={Sun, Lihua and Zhou, Rongxuan and Li, Guiqin and Li, Jianing and Mitrouchev, Peter},
journal={J. Phys.: Conf. Ser.},
volume={2587},
pages={012088},
year={2023}
}
Designed and delivered a production AOI (Automated Optical Inspection) system for Texas Instruments CSE semiconductor products at Dinnar Automation. The system uses 4 industrial CCD cameras (MV-GE501GC) with multi-angle illumination to inspect 19 defect categories including epoxy exposure, pin missing, contamination, and marking defects, achieving 100% detection rate at >85K units/day throughput.
Developed EtherCAT master module with 1 kHz F/T sensor synchronization at CloudMinds Robotics, reducing joint control cycle from 20 ms to 12 ms (40%). Co-designed standardized multi-robot assembly workstations for harmonic reducer production with digital twin systems. Deployed data-driven impedance controller on UR5e with 42% vibration reduction.
Co-designed the full-stack "Fines" robotic platform at SJTU IWIN-FINS Lab: hardware architecture (steerable wheeled chassis), embedded control framework (FineMote, 1 ms cycle), and perception pipeline (FineVision with LiDAR + IMU fusion). The platform enabled experimental validation for the IROS 2025 paper on distributed odometry.
Built GAN-based CAN FD signal synthesizer (5× augmented data, +35% ECU test coverage) and gradient reversal domain adaptation pipeline (CARLA→NIO ET5) at NIO Autonomous Driving, reducing lateral control error by 37% in production vehicle road tests.
Open-source 65-tool MCP server bridging MuJoCo physics simulation with AI assistants via natural language. Supports trajectory optimization (iLQR, MPPI), inverse kinematics, domain randomization, and RL environment integration for 50+ robot models from MuJoCo Menagerie.
Benchmarked FAST-LIO vs. LIO-SAM on real-world driving datasets collected with Northeastern's NUance autonomous vehicle platform. Integrated FAST-LIO-LC for loop closure drift correction over long trajectories, achieving consistent sub-meter accuracy.