Huijuanis our in-house AI simulation engine for autonomous driving, robotics and intelligent hardware development. It covers scene and task orchestration, data generation, quality management and delivery, and it can be adapted to simulation infrastructure a customer already runs.
This page focuses on thesynthetic data generation (SDG) module. The results shown here come from an Unreal Engine 5 implementation, which illustrates the platform's multi-sensor capture, automatic annotation and production-grade data delivery. UE5 is the vehicle for this demonstration, not the only backend the system supports.
Live demo:Huijuan Simulation Platform
Results
SDG synthetic data generation walkthrough
Scene construction, sensor configuration, task deployment and result distribution happen inside a single workflow, which is what makes end-to-end data delivery practical rather than a sequence of manual handoffs.
The platform
SDG data capability at a glance
Where it sits
| Solution | Primary host | Capability profile |
|---|---|---|
| Huijuan SDG data module | In-house architecture; this page shows the UE5 implementation | Data generation, automatic annotation, quality inspection and delivery in one flow, adaptable to the customer's technical environment |
| Isaac Sim | Omniverse | Mature industrial simulation and robotics ecosystem, with a complete sensor suite |
| CARLA | Unreal Engine | Broad open-source autonomous driving ecosystem, convenient for research and algorithm validation |
| AirSim / Colosseum | Unreal Engine | An open solution for vehicle, UAV and robot simulation |
| BlenderProc | Blender | Strong at offline synthetic data generation and high-quality rendering |
Where the differences are (capability matrix)
How the marks read: ★★★ deep coverage; ★★ complete coverage; ★ basic coverage or needs an extension; — not presented as a native capability in public material. The stars indicate breadth of coverage, not a ranking of suitability for any particular project.
| Capability | Huijuan SDG module | Isaac Sim | CARLA | AirSim | BlenderProc |
|---|---|---|---|---|---|
| Real-time data generation | ★★ | ★★ | ★★ | ★★ | — |
| Pinhole, fisheye and 360° panoramic | ★★★ | ★★ | — | — | ★★ |
| Thermal infrared data | ★★★ | ★ | — | ★ | ★ |
| Semantic and instance labels | ★★ | ★★ | ★★ | ★ | ★★ |
| Material-level labels | ★★★ | — | — | — | — |
| Object motion and whole-frame motion labels | ★★ | ★★ | ★★ | — | ★★ |
| Edge structure data | ★★ | ★★ | — | — | — |
| On-vehicle 3D occupancy ground truth | ★★★ | ★ | — | — | — |
| LiDAR scan timing compensation | ★★★ | — | — | — | — |
| Satellite positioning error and occlusion | ★★★ | — | ★ | ★ | — |
| Synchronised camera and inertial export | ★★★ | — | — | — | — |
| Pre-capture quality inspection | ★★★ | — | — | — | — |
| Structured performance report | ★★★ | ★★ | ★ | — | — |
Data output comparison
| Capability | Huijuan (the module shown here) | Isaac Sim | CARLA | AirSim | BlenderProc |
|---|---|---|---|---|---|
| Cameras and panoramas | Pinhole, two fisheye families and 360° panoramic under one configuration | Several camera models | Mostly pinhole | Mostly pinhole | Fisheye supported, offline oriented |
| Thermal infrared | Temperature and thermal radiation data aligned to the visible image | Achievable through an extension | Needs an extension | Customisable | Customisable |
| Distance and geometry | Two distance types, surface orientation and supplementary transparent-object data in the same frame | Fairly complete geometry data | Basic distance data | Basic distance data | High offline accuracy |
| Per-pixel labels | Semantic, instance and material labels | Semantic and instance labels | Semantic and instance labels | Basic labels | Semantic and instance labels |
| Motion labels | Describes both object motion and whole-frame motion | Whole-frame motion supported | Frame motion supported | Needs an extension | Can be generated offline |
| Surface colour reference | Lighting-independent surface colour, which helps cross-domain analysis | Supported | Needs an extension | Needs an extension | Supported |
| Edge structure | Output in the same frame as everything else | Supported via post-processing | Needs an extension | Needs an extension | Can be generated offline |
Sensor and data engineering comparison
| Capability | Huijuan (the module shown here) | Isaac Sim | CARLA | AirSim | BlenderProc |
|---|---|---|---|---|---|
| LiDAR motion compensation | Corrects for motion across the scan | Extended per project | Needs an extension | Needs an extension | — |
| Transparent object handling | Detected automatically and handled by rule | Depends on scene and material setup | Depends on scene setup | Depends on scene setup | — |
| 3D occupancy ground truth | Single-frame, accumulated and moving scans | Oriented to robot navigation | Possible with offline tooling | Needs an extension | — |
| 3D object annotation | Position, pose, visibility relations and projection results | Supported | Basic annotation supported | Needs an extension | Supported |
| Sensor perturbation | Covers camera, LiDAR, inertial and mounting vibration | Fairly complete for cameras and active sensors | Covers some sensors | Covers some sensors | Image post-processing oriented |
| Satellite positioning simulation | Accounts for constellation, occlusion, reflection and the receiving process | Needs an extension | Basic error simulation | Basic error simulation | — |
| Camera and inertial synchronisation | Timing, calibration and data organisation handled inside one task | Organised by the user | Organised by the user | Organised by the user | — |
| Label taxonomy | Switchable between common dataset class standards | Custom taxonomies supported | Mostly preset classes | Custom taxonomies supported | Custom taxonomies supported |
| Output and conversion | Covers common image, array, point cloud and standard annotation formats | Covers the common formats | Covers the common formats | Covers the common formats | Covers the common formats |
| Pre-capture inspection | Scene, sensors, hardware and configuration checked in one place | Standard toolchain checks | Standard runtime checks | Standard runtime checks | Script checks |
Throughput and storage on the same hardware
The numbers below compare the data path before and after optimisation on identical hardware, resolution and task configuration, to show what the module on this page actually gains.
After the colour-image and surface-colour capture paths were parallelised, on a singleRTX 3090at 1920×1080 the critical-path time per frame fell from26.9 ms to 4.32 ms (6.23× faster), and the bytes written per frame dropped by48% (mathematically lossless). Throughput and storage on the same hardware over the same period follow directly:
| Metric (RTX 3090 · 1080p) | Before | Ours (after) | Gain |
|---|---|---|---|
| Critical-path time per frame | 26.9 ms | 4.32 ms | 6.23× faster |
| Throughput per GPU | ≈37.2 frames/s | ≈231.5 frames/s | ≈6.2× |
| Continuous capture over a full day | ≈3.21M frames/day | ≈20.02M frames/day | ≈6.2× |
| Bytes written per frame | Baseline | −48% | storage cost nearly halved |
How this is derived: throughput = 1000 ÷ time per frame (26.9 ms → 37.2 fps, 4.32 ms → 231.5 fps); daily output = throughput × 86,400 seconds (37.2 × 86,400 ≈ 3.21×10⁶, 231.5 × 86,400 ≈ 2.00×10⁷).The conclusion: on the same GPU over the same capture window, roughly 6.2× the output at roughly half the storage per unit of data— which directly compresses the GPU-hours and the storage bill of long-tail data production.
Eight core data capabilities
1) Multiple camera models and panoramic capture
Pinhole cameras, fisheye cameras of differing fields of view and 360° panoramic capture are all supported, with camera parameters, data types and output cadence managed together inside one task. The capture conventions and data interfaces adapt to the simulation environment the customer already runs.
2) Multi-modal output and the "omit rather than mislead" rule
A single capture can emit visible-light imagery, distance, surface orientation, per-pixel labels, motion information, edges and thermal infrared, all aligned. Where a data type does not lend itself to panoramic representation, the system switches that output off rather than producing training data that is structurally complete but semantically wrong.
Thermal infrared is generated from ambient temperature, material thermal properties and sensor response, so the output carries temperature information you can analyse — not a false-colour filter applied to an ordinary image.
3) Distance, label and motion ground truth
- Distance data: describes both the distance along the camera axis and the straight-line distance from object to camera, which suits different perception tasks; transparent objects keep image and distance consistent through supplementary data.
- Per-pixel labels: class, object and material labels at three levels, which trains recognition models and also supports analysis of the domain gap different materials introduce.
- Motion information: object motion and whole-frame change including camera motion are described separately, which suits training for tracking, motion estimation and temporal understanding.
4) LiDAR data closer to a real sensor
A real LiDAR takes time to complete one revolution, and the vehicle and the environment keep moving during it. The module corrects for motion across the scan, so the point cloud retains temporal characteristics closer to a real device; transparent materials such as glass are handled automatically, which removes the need to edit scene configuration object by object.
5) Millimetre-wave radar data
Radar shares one scene and one timeline with the cameras and the label data, and generates range, velocity and bearing from object material, distance and motion state. The exact computation path can be chosen to match the customer's engine and hardware; what matters is that radar, imagery and scene ground truth stay consistent with one another.
6) 3D occupancy ground truth
The system can divide three-dimensional space into a grid and mark whether each cell is occupied, supporting single-frame snapshots, global accumulation and incremental scanning as the platform moves. The data feeds occupancy perception, traversability analysis and spatial planning models directly.
7) Sensor domain randomisation: perturbation at the physical level
Exposure, colour temperature, measurement error, dropped points, device vibration and mounting offset can all be perturbed under control, and a fixed random seed reproduces the same set of variations. That both widens the training distribution and lets different algorithm versions be compared fairly under identical conditions.
8) A production-grade data pipeline: from "we can capture it" to "we can ship it"
- Trustworthy ground truth: 3D object position, pose and image projection all come from the same scene state, and visual spot checks plus consistency checks reduce label error.
- Automatic semantic alignment: manual labels, scene rules and a local semantic model work together, and the class standard of a common dataset can be switched in.
- Formats and fault tolerance: covers common image, array, point cloud and standard annotation formats; scene, sensors, hardware and task configuration are checked together before capture, so a batch does not have to run before unusable data is discovered.
- Capture performance: after the colour-image and surface-colour paths were parallelised, critical-path time per frame at 1920×1080 on an RTX 3090 fell from 26.9 ms to4.32 ms (6.23× faster)and bytes written fell48%, mathematically lossless.
Platform modules
1. Asset library
One place to manage and search assets: people, furniture, vehicles, buildings, materials, appliances, decor, kitchen, bathroom and outdoor across 10 top-level categories and more than 90 intelligent types, with grid and list views, search and filtering, bulk operations and demo data import.
2. Asset groups and distribution
Related assets are organised into reusable groups, and rules control count, position, density and probability of appearance, which produces batches of scenes that differ from one another yet remain traceable.
3. The SDG workflow (five stages)
| Stage | What happens |
|---|---|
| Data preparation | Upload datasets, pick existing data and scene assets |
| Environment setup | Set sensors, scene parameters, sampling strategy and perturbation ranges |
| Task execution | Allocate resources and monitor the run |
| Data processing | Augmentation, inspection and format conversion |
| Delivery | Manage downloads, versions and delivery records |
4. Capture configuration
- Sensor configuration: cameras, fisheye, panoramic, thermal infrared, LiDAR, millimetre-wave radar, satellite positioning, inertial and contact sources
- Capture configuration: total frame count, sampling strategy (uniform / keyframe / intelligent / random), storage mode (live streaming / local / cloud / hybrid)
- Path navigation: straight line / spline / free path / follow / random walk
- General: custom CAD models can be uploaded (OBJ / FBX / GLTF / STL / STEP and others)
5. Scene generation and real-environment reconstruction
- Rule-based scene generation: batches of controlled scenes from templates, constraints and distribution rules
- Natural-language scene description: describe the environment, the objects and the task goal in business language
- Real-environment reconstruction: turn site imagery into spatial assets that can enter the simulation flow
6. Data production walkthrough
A complex data task is broken into steps that can be shown and reproduced, covering data intake, task preparation, environment setup, execution, processing and delivery of results — suitable for customer review, internal training and solution validation.
Data production walkthrough
Production-grade delivery
Huijuan's in-house task engine, data modules, workflows and scheduling are delivered to adeployable-at-scaleengineering standard, and can be adapted to the UE, Omniverse or other simulation infrastructure a customer already runs. The system supports batch tasks, run monitoring, data storage and versioned distribution across local, cloud and hybrid environments, which makes it straightforward to grow from a single-machine trial to a continuously running data production line.
Where it applies
- Autonomous driving: long-tail scene data production, multi-sensor data and annotation generation, scaling out training data
- Robotics: task scene construction for embodied intelligence, multi-modal sensor capture and training data generation
- Game development and virtual production: procedural scene construction, asset reuse and efficient batch content production
Further reading: the SDG series
Fuller implementation detail, validation methods and references are collected in the SDG series: camera and panoramic capture, distance and labels, motion information, active sensors, scene understanding and data workflows.
