Using AI Models in Unreal Engine: Basic Pitch, ONNX and NNE

Table of Contents

Using an AI model in Unreal Engine can feel like a specialist task. But a useful first integration starts with familiar development work: prepare some data, call a system, and turn its output into something your project understands.

For my Music Transcriber, the goal was to turn recorded audio into a Music Session for MidiEngine. Spotify’s Basic Pitch supplied the note predictions. Unreal’s Neural Network Engine (NNE) supplied the interface for running the model. The interesting part was connecting those pieces.

1. Understand the three pieces

Basic Pitch is Spotify’s pretrained model for music transcription. It can detect overlapping notes and works across different instruments, though Spotify recommends one instrument at a time for the best results. That makes an isolated piano or guitar recording a more useful starting point than a busy finished mix.

ONNX is a format for representing the model so a compatible runtime can execute it. NNE is Unreal’s common interface to neural network runtimes. In this integration, I use the Basic Pitch ONNX export through the ONNX Runtime plugin, selecting its CPU runtime, NNERuntimeORTCpu.

This is inference: using a model that has already been trained. The Unreal application prepares inputs and interprets predictions; it does not need to train Basic Pitch or embed the original Python application.

2. Bring the model into Unreal

The setup starts by enabling the appropriate NNE runtime plugin, adding NNE to the C++ module dependencies, and importing the ONNX file as an NNE Model Data asset. My Music Transcriber settings asset holds the reference to the Basic Pitch model.

The code loads that asset, resolves the CPU runtime, creates a model, and then creates a model instance to execute it. Think of the asset as the stored model and the instance as the working session with its execution state and buffers.

An ONNX file is a starting point, not a guarantee of compatibility. The selected runtime must support the model and target platform. Check that early, and make sure the model asset is included when packaging. Epic’s NNE overview explains the asset, runtime and instance relationship.

3. Give the model the input it expects

A model expects numbers arranged in a particular way. These arrays are called tensors; their shape describes their dimensions. Matching that contract is just as important as loading the model successfully.

For Basic Pitch, my backend decodes the SoundWave into audio samples, averages the channels to mono, and resamples to 22,050 Hz. It then supplies overlapping windows of roughly two seconds. The supported export expects each input window as a float tensor shaped [1, 43844, 1].

The overlap gives the model context around window boundaries. After inference, the implementation trims the output edges and joins the retained frames before decoding notes. This avoids treating every window as an unrelated piece of music.

The same principle applies elsewhere: an image model might need resized pixels and a particular channel order; a motion model might need joint positions in a specific coordinate system. Follow the model’s reference preprocessing before inventing your own.

4. Turn predictions into usable data

The backend validates tensor names, types and shapes, binds input and output buffers, and calls NNE’s RunSync. In this implementation, model-instance creation, audio processing and inference happen on a worker thread. The game thread receives the completed result.

Basic Pitch returns predictions for note activity, note starts, and finer pitch contours. These are numerical predictions, not ready-made Unreal note objects. My decoder uses onset and confidence thresholds, pitch limits and minimum note lengths to turn them into note events.

This version produces semitone notes without pitch-bend automation. The output BPM is supplied by the caller; this transcription path does not estimate tempo. Those boundaries matter when deciding what the result can do.

5. Connect the result to the actual feature

The Music Transcriber converts detected note times into beats using that chosen BPM. On the game thread, it builds a Music Session containing a MIDI Pattern, a track, and its notes. The SoundWave remains referenced as background audio. An editor action can save the completed session as a native asset, without exporting and reimporting a MIDI file.

The complete flow
SoundWave → prepared audio → Basic Pitch through NNE → decoded notes → MidiEngine Music Session

In the editor, the integration exposes Transcribe with Basic Pitch… on a selected SoundWave. Runtime Blueprints use Transcribe SoundWave with Basic Pitch and receive the result asynchronously. These are features of the Music Transcriber integration described here.

That final connection was the reason for building it. Once predictions become normal musical data, they can feed playback and music-driven interactions. MidiEngine provides that next layer for projects such as rhythm games and music visualizers, so transcription has a useful destination inside Unreal.

Start small, then apply the pattern elsewhere

For a first experiment, choose a pretrained model with a working reference example, use a small known input, and compare the result before adding your game’s behavior. Keep input preparation, inference and output interpretation separate so you can see where a mismatch begins.

Your next model might classify an image, recognize a gesture or analyze a sound. The model-specific details will change, but the integration questions remain approachable: what does it expect, what does it return, and how will your Unreal project use that result?