Overhead Crane AI Vision Inspection: 5 Core Uses from Safety to SLAM
AI Vision & Detection for Overhead Cranes: 5 Core Applications from Safety Monitoring to SLAM Mapping. AI-powered vision and detection technology is transforming how heavy-industry workshops operate—from safety protection and load positioning to equipment health monitoring and environment mapping. Deep-learning vision now covers every critical aspect of overhead crane intelligence. Kelude Heavy Industry has deployed five core applications built on the NVIDIA Jetson Orin NX edge computing platform with the YOLOv8 framework.
AI vision and detection technology for overhead cranes (bridge cranes) is reshaping operations across heavy-industry workshops—from safety protection and load positioning to equipment health monitoring and environment mapping. Deep-learning vision now spans every critical link in overhead crane intelligence. Kelude Heavy Industry has systematically built an AI vision technology stack covering the full crane operating cycle, based on the NVIDIA Jetson Orin NX edge computing platform and the YOLOv8 object detection framework. This article serves as the core guide to the “AI Vision & Detection for Overhead Cranes” series, offering a panoramic breakdown of the five key application scenarios—covering technical architecture, deployment strategies, and selection recommendations—to provide a complete technical roadmap for overhead crane intelligence upgrades.
| Serial No. | Application Scenario | Core Algorithm | Accuracy Indicator | Hardware Platform | Related Standard |
|---|---|---|---|---|---|
| 1 | AIVision-based Safety Detection | YOLOv8sPersonnel Detection | m AP≥95% | Jetson Orin NX | GB/T 28264 Safety Monitoring and Management System-2017 |
| 2 | Load Vision Positioning | YOLOv8s + Coordinate Transformation | ±5mm | Jetson Orin NX | JB/T 1306-2008 |
| 3 | Wire Rope AIDetection | YOLOv8s+Efficient Net-B4 | m AP≥93% | Jetson Orin NX | GB/T 5972-2016 |
| 4 | Weld Seam AIDetection | YOLOv8s-weld | m AP≥94% | Jetson AGX Orin | GB/T 3323-2005 |
| 5 | Vision SLAM Mapping and Positioning | ORB-SLAM3/VINS-Mono | ±10cm | Jetson Orin/AGX | ISO 4301 Crane Design Standard-2008 |
The Evolution of AI Vision Technology for Overhead Cranes
In backbone industries such as metallurgy, shipbuilding, heavy machinery, and new-energy equipment manufacturing, the overhead crane is the logistical heart of the workshop floor—directly determining both production efficiency and safety performance. Traditional crane operation relies on the operator's visual judgment, voice communication with ground personnel, and scheduled shutdowns for manual inspection. This model is now facing fundamental challenges amid rising production targets, increasing labor costs, and ever-stricter safety compliance requirements.
The maturity of industrial AI vision—particularly deep-learning-based object detection frameworks—has opened a cost-effective, high-return path to crane intelligence. Unlike conventional sensor solutions such as Gray-code bus positioning systems, laser distance measurement, or encoders, an AI vision system uses industrial cameras as its core sensing element and requires no infrastructure along the crane runway. It delivers personnel safety monitoring, load positioning, equipment defect detection, and environment mapping in a single unified platform. Built on a common edge-computing hardware base (NVIDIA Jetson series) and algorithm framework (YOLOv8 series), the crane AI vision system can be deployed incrementally in a modular, crane-by-crane approach—starting with the most critical safety scenarios and expanding toward a fully intelligent crane ecosystem.
Kelude Heavy Industry's in-house developed crane AI vision and inspection technology covers five core application areas: safety monitoring, load positioning, wire rope inspection, weld inspection, and vision-based SLAM mapping and localization. Together, these five scenarios form a complete capability matrix for crane AI vision across five dimensions: people (safety protection), load (material handling), machine (equipment health), structure (metal structure inspection), and position (environment awareness). Each is examined in depth below.
AI Vision Safety Monitoring: Personnel Intrusion Detection and Anti-Collision for Crane Lifting Zones
Personnel safety during crane lifting operations is the top priority in workshop safety management. Per GB/T 28264-2017 Safety Monitoring and Management System for Lifting Appliances, cranes must be capable of actively detecting and responding to hazardous conditions in the lifting zone. Conventional safety measures—physical barriers, safety door locks, and manual supervision—rely on passive isolation and human attention, leaving significant protection gaps in dynamic, complex lifting environments.
Kelude Heavy Industry's AI vision safety monitoring system for overhead cranes deploys one industrial AI camera on each end carriage and one beneath the trolley, providing triple-redundant coverage of the lifting zone. The cameras capture 1080P real-time images at 30fps, transmitted via Gigabit Ethernet to a Jetson Orin NX edge-computing module (100 TOPS). A YOLOv8s object detection model running on the Jetson platform performs person detection on every frame, with single-frame inference latency below 15ms and end-to-end response time under 100ms.
The system's core safety logic is a three-tier alarm strategy—a yellow warning zone (5–8m) triggers an audible and visual alarm plus deceleration; an orange caution zone (3–5m) triggers emergency deceleration and voice-based personnel evacuation; and a red danger zone (<3m) triggers STO (Safe Torque Off) emergency stop. Under normal lighting (500 Lux), person detection achieves mAP ≥95%; under low-light conditions (50 Lux with infrared illumination), mAP ≥90%. The three-camera redundant deployment ensures every ground worker is covered by at least two cameras simultaneously, eliminating missed detections caused by single-view occlusion. When multiple cranes operate on the same runway, a time-division, frequency-offset illumination scheme keeps the false-trigger rate below 0.1 events per hour across three cranes. See AI Vision Safety Monitoring for Overhead Cranes: Engineering Practice in Personnel Intrusion Detection and Anti-Collision for details.
Vision-Based Load Recognition and Precision Positioning for Crane Hoisting Operations
Load positioning—precisely placing steel coils, steel plates, large equipment, or molds at a designated location—is a critical bottleneck in workshop production efficiency. Traditional practice relies on the operator's visual estimation coordinated with ground guidance, and positioning accuracy is limited by skill level and sightlines, often requiring multiple adjustment cycles for a single placement. With the growing adoption of continuous casting and rolling processes in the steel industry and the precision assembly demands of new-energy equipment manufacturing, manual positioning can no longer keep pace with modern production rhythms.
The vision-based load recognition and precision positioning system employs a four-layer architecture (perception—inference—transformation—control) to create a complete closed loop from image capture to automatic positioning. The perception layer uses Hikvision or Basler industrial cameras (5–20MP, GigE Vision interface) with LED strobe illumination to compensate for workshop lighting variations. The inference layer runs a YOLOv8s model that detects four load types in real time—steel coil, steel plate, equipment, and mold—achieving mAP@0.5 of 0.962 with single-frame inference of 12–18ms (TensorRT FP16 optimized).
Positioning accuracy is achieved in the coordinate transformation layer: using pre-calibrated camera intrinsic matrices and perspective transformation matrices, the pixel coordinates output by YOLOv8 are converted to ground-plane world coordinates (in mm). Bilinear interpolation and sub-pixel optimization deliver positioning accuracy of ±5mm. The control layer sends deviation values to a Siemens S7-1200/1500 PLC via a PROFINET gateway, generating motion commands for the crane bridge, trolley, and hoisting mechanism. End-to-end processing latency is kept within 100ms, meeting the real-time control requirements of low-speed crane operation (≤20m/min). See Vision-Based Load Recognition and Precision Positioning for Overhead Cranes: YOLOv8 Load Detection and Vision-Guided Positioning in Engineering Practice.
AI-Powered Online Wire Rope Inspection System for Overhead Cranes
The wire rope is the most critical load-bearing component of an overhead crane, and its condition directly affects equipment and personnel safety. Per GB/T 5972-2016 Cranes—Wire Ropes—Care, Maintenance, Inspection and Discard and TSG Q7015-2016 Rules for Periodic Inspection of Lifting Appliances, wire ropes must be regularly inspected for broken wires, wear, corrosion, and deformation. Traditional manual Visual Testing (VT) depends heavily on inspector experience and suffers from blind spots, inconsistent criteria, and an inability to provide real-time warnings.
The AI vision online wire rope inspection system uses a multi-camera coordinated deployment—a top-down camera covers the fixed-end section of the upper rope, a side camera monitors the free rope length in the middle, and a bottom-up camera targets the lower rope end and the hook block sheave area. The perception layer uses global-shutter CMOS sensors (up to 200fps) with LED strobe illumination whose pulse width can be adjusted down to the microsecond level. The inference layer runs an improved YOLOv8s model (incorporating DCNv4 deformable convolutions and Coordinate Attention) to detect three primary defect types—broken wires, wear, and corrosion—in real time, achieving a combined detection mAP ≥0.93.
At the analysis layer, the system uses an EfficientNet-B4 backbone network to perform secondary fine-grained classification of 11 defect subtypes (localized broken wires, concentrated broken wires, uniform wear, localized wear, pitting corrosion, uniform corrosion, wavy deformation, lantern deformation, rope diameter reduction, poor lubrication, and surface foreign matter adhesion), achieving a classification accuracy of ≥95%. Defects are graded on a severity scale from Grade 0 to Grade 3, with automatic alarms triggered at moderate severity and above. The alarm layer integrates with PHM (Predictive Health Management) and CMMS (Computerized Maintenance Management System), reducing the average response time from detection to work order generation to no more than 30 seconds. The multispectral fusion strategy (RGB+NIR) improves corrosion defect detection recall by 12.7%. For further details, refer to "AI Vision Online Inspection System for Crane Wire Rope: Engineering Practice in Deep Learning-Based Broken Wire, Wear, and Corrosion Defect Identification".
5. AI Vision Inspection System for Crane Rail and Main Girder Weld Seams
The butt welds on crane rails and the fillet welds on main girders represent the most vulnerable points in the entire metal structure of an overhead crane. These welds endure long-term alternating loads, heavy-impact shocks, and environmental corrosion—and if cracks or incomplete penetration defects go undetected, the consequences can be catastrophic structural failure. AI vision inspection of weld defects presents significant technical challenges: crack widths of only 0.05–0.3 mm make this a classic small-target detection problem, while weld surfaces are contaminated with strong interference backgrounds such as slag, spatter, and oxide scale.
The AI Vision Inspection System for Crane Rail and Main Girder Weld Seams adopts a three-form deployment architecture—an AGV autonomous inspection vehicle (for automated large-area rail weld inspection), a handheld detector (for confined spaces and end carriage fillet welds), and a fixed work station (at the main girder assembly and welding station)—with all three forms sharing a unified software stack. The perception layer is equipped with dual industrial cameras (global shutter wide-angle + macro), a structured-light ring illuminator, and a laser profile scanner (to capture the three-dimensional weld profile).
The inference layer runs the YOLOv8s-weld model—pruned and optimized from YOLOv8s with an added small-target detection head specifically designed for weld defect features—achieving single-frame inference of ≤15 ms (≥60 FPS) on a Jetson AGX Orin. The system provides real-time identification of five defect categories: cracks (longitudinal, transverse, and crater subtypes), incomplete penetration, porosity, slag inclusion, and undercut, with a combined mAP of ≥0.94. The evaluation layer automatically aligns with the GB/T 3323-2020 quality grading standard for radiographic testing of fusion-welded joints, outputting Grade I–IV quality ratings along with rework recommendations. Weld types covered include QU70/80/100 rail butt welds with V-groove preparation, double-sided fillet welds on main girders, and mixed welds on end carriages—the three typical weld configurations found on overhead cranes. For further details, refer to "AI Vision Inspection System for Crane Rail and Main Girder Weld Seams: Engineering Practice in Deep Learning-Based Crack, Incomplete Penetration, and Porosity Defect Identification".
6. Visual SLAM Environment Mapping and Positioning System for Overhead Cranes
One of the core bottlenecks in fully automatic overhead crane operation is achieving highly reliable, high-precision global positioning inside factory buildings where GPS is unavailable. Traditional approaches—laser distance sensors that suffer severely from dust obstruction, Gray-code bus positioning systems with high installation and maintenance costs, and encoders that accumulate drift from wheel slippage—can only provide one- or two-dimensional position information and cannot perceive the crane's six-degree-of-freedom spatial orientation.
The Visual SLAM Environment Mapping and Positioning System for Overhead Cranes integrates two open-source frameworks, ORB-SLAM3 and VINS-Mono, using monocular/stereo cameras plus a six-axis IMU as sensors to simultaneously perform six-degree-of-freedom pose estimation and sparse/dense map construction in unknown factory environments. Factory feature extraction uses crane-adapted parameters (nFeatures increased from 1,200 to 2,000, iniThFAST reduced from 20 to 12), with an LSD line-segment feature layer added to assist the matching stage, improving the feature matching rate from 35% to 52%.
The IMU pre-integration combined with tight visual-inertial coupling reduces the tracking loss rate under vibration conditions from approximately 8% in vision-only mode to 0.5%. To address map scale expansion in large-span factory buildings (100–300 m), the system employs three optimization measures: covisibility graph pruning (triggered every 50 frames, removing 5–10% of redundant data), active local submap management (50 m × 30 m grid cell segmentation), and incremental loop closure detection (DBoW2 bag-of-words plus pose graph optimization, completing 500 ms for a 3,000-frame map), achieving positioning accuracy of ±10 cm. VINS-Mono's sliding-window nonlinear optimization approach offers greater real-time advantages in edge deployment scenarios with constrained computing resources. This system can be deployed independently as a global positioning redundancy for the crane, or deeply fused with laser distance measurement and encoders to form a multi-sensor positioning architecture. For further details, refer to "Visual SLAM Environment Mapping and Positioning System for Overhead Cranes".
7. Technical Comparison and Deployment Selection Guide for the Five Systems
The five systems each have distinct strengths in technology stack, performance metrics, and hardware costs. In practical deployment, selection should be tailored to the specific crane operating conditions, workshop environment, and maintenance management requirements. The table below provides a horizontal comparison across key parameter dimensions:
| Comparison Parameter | Vision-based Safety Detection | Load Positioning | Wire Rope Inspection | Weld Inspection | Load Vision SLAM |
|---|---|---|---|---|---|
| Detection Object | person(Personnel) | Hook_Load(Load) | Broken Wire/Wear/Corrosion | Crack/Porosity/Incomplete Penetration | Environmental Feature Points/IMU |
| Inference Latency | ≤15ms | 12~18ms | ≤25ms | ≤15ms | 33~50ms |
| Accuracy/Indicator | m AP≥95% | ±5mm | m AP≥93% | m AP≥94% | ±10cm |
| Number of Cameras | 2~4Hardware Platform | 1~2Hardware Platform | 3~6Hardware Platform | 2Hardware Platform+Laserobturator Instrument | 1~2Hardware Platform+IMU |
| Recommended Deployment | Alloverhead crane(Primary Recommendation) | Precise Positioning Scenario | Highsafety leveloverhead crane | Periodic Inspection/Newly Manufactured Inspection | Fully automaticoverhead crane/redundancy Positioning |
Deployment priority recommendations: For retrofitting a single overhead crane, we recommend a phased approach that follows a "safety → health → positioning → mapping" progression. The personnel intrusion detection system is the foundational safety layer and should be the first mandatory deployment on every crane. Wire Rope Inspection and Weld Inspection can be selectively deployed based on the crane's service age and operating environment — older cranes should prioritize Wire Rope Inspection, while newly manufactured cranes should prioritize factory weld inspection. The load positioning system addresses production cycle bottlenecks and suits cranes with automated alignment requirements. Vision-based SLAM is designed for fully automatic operation and smart logistics scenarios. All five systems share the Jetson Orin NX/AGX Edge Computing platform and the YOLOv8 algorithm ecosystem, enabling integrated deployment on a single crane with resource sharing through GPU time-slice allocation and dynamic model loading.
Multi-Sensor Fusion and System Integration Outlook
Sensor data from the five systems is deeply fused on a unified Edge Computing platform, forming a closed perception-decision-control loop that covers the full spectrum of crane operations. Take the coordinated response between safety monitoring and load positioning as an example: when the personnel intrusion detection system identifies a worker entering the orange warning zone, the load positioning system simultaneously records the current load coordinates, providing complete spatial-temporal data for subsequent safety incident traceability. When the Wire Rope Inspection system detects a Grade 3 critical defect, it automatically triggers a focused inspection task in the Weld Inspection system — the weld inspection AGV receives the command, autonomously plans a route to the crane rail segment near the defect location, and performs targeted inspection of the rail welds and Main Girder fillet welds in that area.
From a broader perspective, these five AI vision systems integrate with the crane's core control system (PLC), Predictive Health Management platform (PHM), Manufacturing Execution System (MES), and Enterprise Resource Planning (ERP) system to form a complete Industrial Internet of Things (IoT) architecture. Real-time data collected at the perception layer is processed through edge AI, with structured results uploaded to workshop-level and enterprise-level management platforms. This provides the data foundation for full life cycle management, production scheduling optimization, and safety compliance auditing. As edge AI computing power continues to advance (the next-generation Jetson Orin platform is expected to deliver 200–500 TOPS) and multimodal large models evolve, crane AI vision systems will progress from single-purpose detection to cognitive understanding — not only detecting personnel intrusion but also interpreting hand gestures and work intent of workers on the ground.