ISAR

Immersive Audio for Split Rendering Scenarios

Services →
Introduced in Rel-18

ISAR is a 3GPP framework for delivering high-quality, low-latency spatial audio in Extended Reality applications where the audio rendering is split between a device and the network.

Category
Services
Introduced
Rel-18
Where
Services › Codecs
Specifications
7 specs
ISAR Description Purpose Related Classification Detected Changes Specifications

Description

Immersive Audio for Split Rendering Scenarios (ISAR) is a 3GPP media codec and system specification designed to deliver high-quality, object-based spatial audio for Extended Reality (XR) applications, particularly in network architectures where rendering is split between a user device (e.g., XR headset) and a network edge server. ISAR addresses the unique challenges of streaming immersive audio, which requires six-degrees-of-freedom (6DoF) rendering, low end-to-end latency, and high compression efficiency to conserve bandwidth. The architecture typically involves an XR application server in the network (e.g., at the edge) that generates or processes the raw audio scene, containing multiple audio objects with metadata describing their positions, orientations, and acoustic properties. The ISAR encoder compresses this audio scene. A key innovation is the split of the rendering pipeline: part of the rendering (e.g., early reflections, basic binauralization) can be performed on the server, while the final stage (e.g., late reverberation, personalized head-related transfer function (HRTF) application, and compensation for last-moment head movements) is performed on the user equipment (UE). This split reduces the data rate that needs to be transmitted compared to sending fully rendered binaural audio, while also offloading complex processing from the potentially resource-constrained UE. The ISAR stream, containing encoded audio objects and rendering metadata, is delivered over the 5G network. The UE's ISAR decoder and renderer then complete the audio rendering based on the latest sensor data (head position) to create a precise, personalized spatial audio experience. The specifications (TS 26.249, 26.251, etc.) define the codec formats, metadata schemas, APIs, and system interfaces to enable this interoperable, low-latency immersive audio service.

Purpose & Motivation

ISAR was created to solve the audio delivery challenges for truly immersive and interactive XR experiences over mobile networks. Traditional audio codecs (like MPEG-H 3D Audio or Dolby Atmos) are designed for cinematic or broadcast scenarios with fixed playback environments and higher latency tolerance. For interactive XR, where a user can move their head and body in real-time, audio must be rendered dynamically with ultra-low latency (<20ms) to match the visual scene and prevent motion sickness. Transmitting fully rendered binaural audio for every possible head position is prohibitively bandwidth-intensive. ISAR's purpose is to enable efficient streaming by adopting a split-rendering model, which aligns with the overall XR split rendering paradigm studied in 3GPP. This model leverages the compute resources of the 5G network edge for heavy audio processing while keeping final, user-specific rendering on the device. It addresses the limitations of previous approaches: either high bandwidth consumption (sending pre-rendered audio) or high device compute load (rendering everything locally from raw objects, which may not be feasible on lightweight XR glasses). By standardizing ISAR, 3GPP aims to ensure interoperability between XR application providers, network operators, and device manufacturers, fostering a ecosystem for high-quality cloud/edge-rendered XR services over 5G and beyond.

Classification

Part ofXR
Related approaches6DOF

Release Timeline

Detected Changes Across Releases

from 3GPP Change Requests

Specific changes extracted from the „Change history“ tables of 3GPP specifications (1 CRs across 1 releases). Complements the general historical overview above with the evidence-based evolution of this function.

Rel-18 1 change

In Release 18, the ISAR (Immersive Audio for Split Rendering Scenarios) function was newly introduced, providing a detailed algorithmic description for split rendering applicable to immersive audio systems. This specifically added the ISAR track-a split rendering feature, including defined interfaces for both pre-renderer and post-renderer operations, with the IVAS codec's split rendering feature serving as the baseline. The release established mandatory post-renderer procedures for compliant UEs and enabled operation at various Degrees of Freedom (DOF) for pose correction.

  • Adding ISAR track-a split rendering feature to TS 26.258 and Corrections to the IVAS C-Code and corresponding specification text TS 26.258CR0002

Explore further

Broader topics and technologies where ISAR plays a role.

Defining Specifications

3GPP specifications that define or reference ISAR, with the latest known release. Sourced from the 3GPP document catalog — see methodology.

SpecificationTitleRelease
TS 26.249 vj00 Immersive Audio Split Rendering (ISAR) Rel-19
TS 26.251 vj00 IVAS Codec Fixed-Point C Code Specification Rel-19
TS 26.252 vj00 IVAS Codec Test Sequences Specification Rel-19
TS 26.258 vj10 IVAS Codec Floating-Point C Code Specification Rel-19
TS 26.260 vj00 Immersive Audio Objective Test Methods Rel-19
TR 26.996 vj00 ISAR Split Rendering Audio Characterization Rel-19
TR 26.997 vj00 IVAS Codec Specification Rel-19
Patrick Zandl

About the author: Patrick Zandl (b. 1974)

Telecommunications specialist, technology journalist (founder of the Mobil server), and developer who has been running since 2025 — the largest Czech-language resource on AI-assisted programming. Formerly Chief Wizard Architect at Prusa3D and head of development for Turris at CZ.NIC; currently a consultant and instructor on AI implementation in companies.