Crane Anti-Sway Algorithm Deep Dive: Kelude's PPO to SAC Engineering

Kelude Heavy Industry has built a complete closed-loop anti-sway control system based on PPO and SAC deep reinforcement learning algorithms, covering everything from simulation training to deployment on S7-1500 PLC hardware. The RL-based anti-sway system reduces settling time by 74.4% compared to PID control, with a residual swing angle of only 0.3° — and it requires no parameter retuning when rope length or load conditions change.

Anti-Sway Control in Overhead Cranes: Why Traditional Methods Fall Short and How AI Solves It

Kelude Heavy Industry · RL Anti-Sway Control Architecture
Sensors
Angle · Position · Swing
State Observer
Filtering · Normalization
RL Agent
PPO / SAC Policy Network
Action Commands
Velocity · Torque
Crane System
Bridge · Trolley · Lifting Spreader
Environment Feedback (Reward Signal · Next State) — Closed-Loop Iterative Training
In industrial environments such as ports, steel mills, and logistics warehouses, load swing in bridge cranes is a critical bottleneck that directly impacts both operational efficiency and safety. Residual oscillation occurs every time the crane starts or stops, and operators traditionally rely on experience to manually damp the swing — a single positioning cycle can take tens of seconds or even minutes of repeated adjustment, severely disrupting workflow throughput. Conventional anti-sway methods — mechanical damping, electronic PID control, and input shaping — perform reasonably well under fixed operating conditions, but their robustness degrades significantly when faced with variable rope lengths, changing load masses, or wind disturbances, requiring frequent manual retuning of parameters. Since 2020, reinforcement learning (RL) has opened a fundamentally new path for anti-sway control. RL agents learn optimal control policies through continuous interaction with the environment, eliminating the need for explicit dynamic system modeling while offering powerful nonlinear fitting capability and inherent adaptability. Kelude's engineering team was the first in the smart crane field to bring two deep RL algorithms — PPO (Proximal Policy Optimization) and SAC (Soft Actor-Critic) — into production-grade anti-sway control systems. This article provides a systematic breakdown of the complete technical stack, from algorithm fundamentals and simulation environment setup to training optimization and industrial deployment, along with quantitative comparisons against classical methods.

State Space and Action Space Design for RL Crane Control

RL Anti-sway's primary task is to map the crane's physical system into a Markov Decision Process (MDP), where the definition of state and action spaces directly impacts policy learning effectiveness and convergence speed. The Kelude solution's state space encompasses five core dimensions:
  • Spreader swing angle θ: Measured in real time by an inclination sensor or vision system, with accuracy typically within ±0.1°. This is the most critical feedback variable for anti-sway, directly reflecting swing amplitude.
  • Angular velocity ω: The first derivative of the swing angle, indicating swing trend and direction. It helps the agent predict future state changes and apply damping actions proactively.
  • Trolley position x: Obtained via encoder feedback for precise positioning—the ultimate goal of anti-sway is to bring the trolley to a stable stop directly above the target position with minimal residual swing.
  • Trolley speed v: A speed closed-loop feedback value that prevents secondary swinging caused by abrupt acceleration changes, while also helping the agent understand the system's momentum characteristics.
  • Rope length L: The real-time rope length from the hoisting mechanism, which directly affects the system's natural frequency (T=2π√(L/g)). Longer ropes result in longer swing periods, requiring the controller policy to adapt accordingly.
The dimensionality of the state vector directly determines the input layer size of the policy network. These five dimensions physically cover all observable variables of the underactuated system, ensuring the RL agent has complete system observability. In implementation, each state dimension must be normalized to the [-1, 1] or [0, 1] range to prevent unstable network training caused by differing units or scales. The action space is designed as a two-dimensional continuous output:
  • Trolley acceleration command ax: Range [-1.0, 1.0] m/s², serving as the speed loop target for the frequency controller. Positive acceleration drives the trolley forward, while negative values apply deceleration or braking.
  • Hoisting speed compensation Δvh: Compensates for hoisting speed during rope length changes. When the trolley moves while hoisting simultaneously, rope length variation alters the system's dynamic characteristics; this compensation term prevents additional excitation disturbances caused by rope length changes.
Compared to discrete actions (e.g., only three levels: accelerate/decelerate/constant speed), a continuous action space enables smoother control curves—critical for heavy-load equipment like cranes, where frequent hard starts and stops not only accelerate mechanical fatigue and shorten wire rope service life but also pose safety risks such as load detachment. Two-dimensional continuous action output allows the agent to execute a perfect anti-sway positioning maneuver with smooth acceleration profiles, much like an experienced operator.

Multi-Objective Reward Function Design for Anti-Sway Control

The reward function is the core of RL training—it encodes the engineer's domain knowledge into the agent's learning objectives. The Kelude solution employs a weighted composite reward comprising four sub-terms: R = w₁·Rswing + w₂·Rpos + w₃·Renergy + w₄·Rterminal
  • Swing angle penalty Rswing = -α·θ² – β·ω². This term applies continuous negative reward whenever the swing angle or angular velocity is non-zero, driving the agent to prioritize swing elimination. The α-to-β ratio controls whether the system favors "position priority" or "swing angle priority"—safety-critical scenarios can increase the αβ value, while efficiency-focused scenarios may reduce it.
  • Positioning reward Rpos = -γ·|x – xtarget|. A larger negative reward is given when the trolley is far from the target position, encouraging the agent to reach the destination promptly. Combined with the terminal reward, this prevents the agent from focusing solely on swing elimination without completing positioning.
  • Energy consumption term Renergy = -δ·ax² – ε·Δvh². Penalizes unnecessary acceleration and compensation actions. This design pushes the agent to find a balance between control quality and energy efficiency, avoiding "over-control" behaviors like high-frequency jitter.
  • Terminal reward Rterminal: A one-time positive reward is granted when |θ| < θth (swing angle threshold, typically 0.5°) and |x – xtarget| < dth (positioning threshold, set to 10mm) are maintained for more than 3 seconds, encouraging the agent to complete the task in a stable state.
The weight coefficients {w₁, w₂, w₃, w₄} are determined through hyperparameter scanning. The Kelude team employs Bayesian Optimization to search for the optimal weight combination in the parameter space. The final principle established is: swing angle penalty weight is the largest (w₁ as the first priority), positioning weight comes second (w₂ as the second priority), and energy consumption weight is the smallest (w₃ as a soft constraint). This hierarchical design ensures the agent first learns to eliminate swing, then precise stopping and energy saving—aligning with the primary principle of industrial safety.

Simulation Environment Setup: Simulink + Adams Co-Simulation

RL training requires millions of time steps of interaction data; training directly on physical equipment is neither safe nor economical—a single erroneous action could cause crane overturning or wire rope fracture. The Kelude solution builds a high-fidelity co-simulation environment based on MATLAB/Simulink + Adams, combining Simulink's algorithmic flexibility with Adams' multi-body dynamics accuracy. Simulink hosts a variable rope length pendulum model, whose core dynamic equation is: (mL²)·θ” + 2mL·vh·θ’ + mgL·sinθ = -mL·cosθ·ax Where m is the load mass, L is the rope length, vh is the hoisting speed, and ax is the trolley acceleration. On the left side of the equation, the first term represents inertial torque, the second term represents Coriolis torque (arising from rope length changes), and the third term represents gravitational restoring torque; the right side represents the excitation torque transmitted from trolley acceleration to the load through the suspension rope. Although this model assumes rigid bodies, it captures the primary dynamic characteristics of crane anti-sway. The Adams side provides detailed 3D multi-body dynamics simulation, incorporating nonlinear factors such as flexible wire rope modeling (discretized multi-segment rigid bodies with spring-damper connections), wheel-rail contact friction (Coulomb friction model with Stribeck effect), and motor torque characteristics (including saturation limits and response delays). Simulink exchanges data every 10ms through the co-simulation interface (Simulink Coder / Adams Controls Toolkit)—Simulink sends acceleration commands to Adams, while Adams feeds back swing angles, positions, and other states to Simulink—forming a high-fidelity closed loop. The advantage of this co-simulation approach lies in leveraging each tool's strengths: Simulink handles rapid RL algorithm iteration and hyperparameter tuning, while Adams ensures physical accuracy for validation. The Kelude team discovered in real projects that RL policies trained purely in the Simulink single-pendulum model showed performance degradation of approximately 15% to 20% when tested in the Adams high-fidelity environment, but recovered to near-simulation levels after further fine-tuning with Adams data.

PPO vs. SAC: Training Comparison and Hyperparameter Tuning

Kelude has conducted engineering validation and systematic comparison of both PPO (Proximal Policy Optimization) and SAC (Soft Actor-Critic) algorithms—marking the first time in the domestic crane industry that two mainstream RL algorithms have undergone complete benchmarking on the same platform. PPO is renowned for its training stability and hyperparameter robustness, converging quickly in anti-sway tasks—reaching a usable level with sway elimination time under 3 seconds in approximately 500,000 time steps. Its core mechanism uses a clip operation to limit policy update step sizes, preventing single updates from deviating excessively from the old policy and causing training collapse. Key hyperparameters include: clip epsilon=0.2, learning rate=3×10⁻⁴, GAE λ=0.95, KL divergence target=0.02, and a network architecture of two hidden layers with 256 nodes each. PPO's clip mechanism offers unique advantages in industrial projects: even with imperfect reward function design, PPO converges robustly, reducing the engineer's tuning expertise requirement. SAC excels in exploration efficiency and final performance. Based on the Maximum Entropy framework, SAC maximizes policy entropy alongside reward maximization, encouraging thorough exploration of the action space during training. After approximately 800,000 time steps, SAC approaches a superior solution: average residual swing angle reduced by about 15%, though early training curves fluctuate more than PPO and show greater sensitivity to reward function hyperparameters. Key SAC hyperparameters include: temperature coefficient α=0.2 (with automatic adjustment mechanism), target entropy=-dim(A) (where A is the action space dimension, here 2), and the same two-layer 256-node network structure. Training curve analysis reveals interesting differences: PPO's cumulative reward converges faster but plateaus later, reflecting KL constraints limiting further policy improvement; SAC's curve shows larger early oscillations but maintains an upward trend later, demonstrating stronger exploration stamina. Based on these findings, the Kelude team innovatively adopted a two-stage training strategy—Stage one uses PPO (1 million time steps) to quickly obtain a reliable initial policy; Stage two starts from this policy and switches to SAC (an additional 500,000 time steps) for fine-tuning, ultimately achieving approximately 12% better overall performance than using either algorithm alone.

Sim2Real Transfer: Domain Randomization Strategy

Sim-to-real transfer remains one of the biggest bottlenecks in industrial reinforcement learning (RL) deployment—and the hardest gap to bridge between academic research and real-world engineering. Simulation environments can never perfectly replicate every physical characteristic of the real world: sensor noise, friction anisotropy, motor response latency fluctuations, nonlinear wire rope stiffness, air resistance, and more. Kelude's approach employs a Domain Randomization strategy to systematically close this gap. The implementation works by randomly sampling key physical parameters at the start of each training episode:
  • Rope length L: Uniformly sampled from 3 m to 15 m, covering the most common operating ranges of overhead cranes.
  • Load mass m: Randomized between 500 kg and 10,000 kg to simulate everything from no-load to full-load conditions.
  • Friction coefficient: Varied by ±30% around the nominal value to represent different wear levels and lubrication states.
  • Sensor noise: Gaussian noise N(0, σ²) added to swing-angle measurements, with σ randomly selected between 0.01 and 0.05 rad to mimic electrical interference in industrial environments.
  • Motor response delay: Uniformly sampled from 5 to 20 ms to reflect VFD response time variations.
  • Wire rope damping coefficient: Varied by ±50% around the nominal value to capture damping differences across rope lengths and diameters.
This "thriving under chaos" training approach forces the policy network to learn robust anti-sway strategies that hold up under a wide range of uncertain conditions—much like an athlete who trains in all weather and can perform consistently anywhere. Kelude's real-machine testing confirmed the effectiveness of Domain Randomization: RL policies trained without it achieved a sim-to-real transfer success rate of under 30%, and the first few runs after deployment almost always produced severe oscillations. In contrast, policies trained with thorough Domain Randomization achieved a transfer success rate above 85%, with the very first real-machine run delivering control quality close to simulation levels.

Comparison with Classic Anti-Sway Methods

Kelude's team conducted a systematic head-to-head comparison of three approaches on the same standardized test platform: PID electronic anti-sway (after engineering tuning), input shaping (Zero Vibration Shaper, ZV), and RL-based anti-sway (a two-stage PPO + SAC integrated training scheme). The test conditions were highly standardized to ensure objective, reproducible comparison data. Test parameters: lifting capacity 10 t, rope length 10 m, trolley travel distance 20 m, and target positioning accuracy of ±10 mm. Sway suppression time: RL anti-sway led the field with a 2.1-second suppression time—74.4% faster than PID's 8.2 seconds and 62.5% faster than ZV input shaping's 5.6 seconds. In a single lifting cycle, this translates to saving 3.5 to 6.1 seconds of waiting time for oscillations to settle. At 30 cycles per hour, that's 105 to 183 seconds saved per hour—equivalent to completing 2 to 3 additional lifting operations per hour. Residual swing angle: RL anti-sway achieved a residual swing angle of just 0.3°, far outperforming PID's 1.8° and ZV's 0.9°. A smaller residual angle means the operator can immediately proceed with spreader alignment and hook release after positioning, without waiting for oscillations to decay—significantly shortening the per-cycle operation time. Robustness: In comparative tests where the rope length varied from 6 m to 14 m, PID required manual re-tuning of parameters or its suppression time degraded to over 12 seconds. The ZV input shaper needed its shaping pulse parameters recalculated, otherwise severe secondary oscillations occurred. The RL approach, by contrast, required no adjustments whatsoever—suppression time stayed consistently between 2.1 and 2.8 seconds across the entire rope length range, demonstrating excellent adaptive capability. It's important to note that RL anti-sway doesn't completely replace traditional methods—rather, it builds on top of mature, reliable underlying control. The speed closed loop still uses a PID speed loop, and RL only outputs optimal decision sequences at the acceleration level. This hybrid "RL decision-making + PID execution" architecture combines the intelligence of the policy with the reliability of the underlying control layer, representing the current mainstream paradigm for industrial RL deployment.

Real-World Deployment: ONNX Model on S7-1500 PLC

Deploying a deep neural network onto an industrial PLC is the final piece of the puzzle—and the most demanding engineering challenge. Traditional deep learning inference frameworks (such as PyTorch and TensorFlow) cannot run directly on PLCs, so Kelude uses ONNX Runtime as the cross-platform inference engine. The trained policy network is exported to the standard ONNX format and deployed on a Siemens S7-1500 PLC. Model lightweighting is the technical core of the deployment solution, involving three key steps:
  • Quantization: Static quantization compresses 32-bit floating-point (FP32) weights and activations down to 8-bit integers (INT8), reducing the model size from 2.3 MB to 0.6 MB—a 74% reduction. The quantized model's inference precision loss stays within 3%, which has no perceptible impact on control quality.
  • Layer Fusion: Consecutive BatchNorm + ReLU + Conv (or fully connected) operators are fused into a single computation node, reducing memory bandwidth bottlenecks and improving inference speed by approximately 40%.
  • Network Pruning: Magnitude-based structured pruning reduces the hidden layer from 256 to 128 nodes while removing all connection weights with a contribution below 0.1%. The final model has approximately 56% fewer parameters, with a performance loss of less than 3%.
The final model deployed on the S7-1500 is a single fully connected neural network: an input layer with 5 nodes (θ, ω, x, v, L), a hidden layer with 128 nodes (ReLU activation), and an output layer with 2 nodes (tanh activation, outputting ax and Δvh). Single inference takes approximately 0.8 ms (measured on the S7-1500's local CPU), fully meeting the real-time requirement of a 10 ms control cycle. The PLC calls the ONNX Runtime dynamic library through a custom function block (FB), executing forward inference every 10 ms in a standard cyclic manner and outputting action commands. For safety, the deployment includes output limiting and degraded-mode protection logic: if ONNX Runtime inference fails or outputs exceed preset safety thresholds, the system automatically falls back to the underlying PID control, ensuring the crane never executes dangerous action commands under any circumstances.

Summary and Outlook

Kelude's RL anti-sway algorithm successfully advances both PPO and SAC deep reinforcement learning from academic research into industrial practice. Through meticulous state-space design, multi-objective reward functions, Simulink + Adams co-simulation validation, Domain Randomization transfer strategies, and lightweight ONNX model deployment on the S7-1500 PLC, the solution forms a complete "simulation training → real-machine deployment" technical closed loop. The approach significantly outperforms classic methods like PID and input shaping in sway suppression efficiency, positioning accuracy, and adaptability across operating conditions, cutting sway suppression time by over 60%. It provides a practical, viable technical path toward intelligent and unmanned industrial crane operations. Looking ahead, Kelude will continue exploring more advanced directions—including multi-crane coordinated anti-sway, machine vision-based direct swing-angle observation, edge-side online adaptive learning, and multi-agent coordinated scheduling—to keep driving lifting appliances from automation toward full intelligence.

References

Frequently Asked Questions (FAQ)

Q: Which reinforcement learning algorithm—PPO or SAC—performs better for crane anti-sway control?
A: PPO converges faster (approximately 500,000 time steps) and offers greater training stability, making it ideal for quickly obtaining a viable solution. SAC achieves superior final performance (residual sway angle reduced by roughly 15%) but exhibits more volatile training curves and requires about 800,000 time steps. The Kelude team employs a two-stage strategy: first using PPO for rapid convergence, then fine-tuning with SAC for optimization—delivering approximately 12% better overall performance than using either algorithm alone.
Q: How is the RL anti-sway algorithm transferred from simulation to real-world cranes?
A: Kelude employs a Domain Randomization strategy, where parameters such as rope length (3m to 15m), load mass (500kg to 10,000kg), friction (±30%), sensor noise, and motor delay (5 to 20ms) are randomized during training. By training the policy across diverse environments to learn a robust control strategy, the Sim2Real transfer success rate has been improved from under 30% to over 85%.
Q: How does a deep reinforcement learning model run on an S7-1500 PLC?
A: The trained PPO/SAC policy network is compressed via INT8 quantization (from 2.3MB to 0.6MB), layer fusion (40% faster inference), and network pruning (56% fewer parameters), then deployed as a three-layer fully connected network (5 inputs, 128 hidden units with ReLU, 2 outputs with tanh). Single inference takes approximately 0.8ms, well within the 10ms control cycle of the S7-1500 PLC, and is safeguarded by output limiting and degradation protection mechanisms.
Q: What advantages does RL anti-sway offer over traditional PID and input-shaping methods?
A: RL anti-sway reduces sway suppression time by 74.4% compared to PID (2.1s vs. 8.2s) and by 62.5% versus ZV input shaping (2.1s vs. 5.6s). Residual sway angle is only 0.3° (PID: 1.8°), and positioning error is 8 mm (PID: 25 mm). Moreover, it requires no parameter retuning when rope length or load changes, delivering significantly better adaptability and robustness than classical methods.

Related Standards

GB/T 6974.6-2008 — Cranes — Terminology — Part 6: Railway Cranes
GB/T 6974.7-2008 — Cranes — Terminology — Part 7: Floating Cranes
GB/T 24818.4-2009 — Cranes — Access, Guarding and Restraint Means — Part 4: Jib Cranes

Related News

contact

contact us

phone:
+86 13903802779

mail:3915269@qq.com

Working hours: Monday to Friday

Wechat
Wechat
SHARE
TOP