What an immersive master actually is
It is not a stereo file with an effect on it. It is a multichannel file in a BW64 container — the long-form wave format defined in ITU-R BS.2088, which exists because the original wave format cannot exceed four gigabytes — carrying two extra chunks. The axml chunk holds Audio Definition Model metadata as XML, defined in ITU-R BS.2076. The chna chunk maps each track in the file to an identifier in that metadata. Without both chunks you have an ordinary multichannel wave file, and it will be rejected as an immersive delivery.
The bed and the objects
The published renderer takes up to one hundred and twenty-eight inputs. The first ten are the bed — a fixed channel layout, commonly seven point one point two. Everything from input eleven onward is an object: a mono or stereo element carrying position metadata that moves it through the room rather than assigning it to a speaker. That gives up to one hundred and eighteen object channel paths. There is no published music-specific limit below that; advice to keep music to a bed and a couple of dozen objects is convention from practice, not a rule.
The numbers the services publish
Apple’s asset guide states the integrated loudness of an immersive master should not exceed minus eighteen, measured per ITU-R BS.1770, and that true peak should not exceed minus one decibel true peak. It requires twenty-four-bit linear audio at forty-eight kilohertz, delivered as a BW64 file with ADM metadata. It also requires a conformed stereo version, synchronised to the immersive one. Note the wording: those loudness figures are a ceiling, not a normalisation target, and the difference matters when you are deciding how hard to push.
Upmixing a stereo mix is not allowed
This is the part people are most often sold and most often surprised by. Apple’s published guidance is explicit that immersive files generated from stereo mixes are not allowed. An immersive master is built from the multitrack session, by a person, in a licensed renderer, with decisions made about what belongs in the bed and what becomes an object. Any service offering to convert a finished stereo master into an immersive one is either doing something the specification forbids or describing something else.
Why binaural is what most people will actually hear
Very few listeners have a speaker array. The overwhelming majority hear immersive music on headphones, through a binaural render. The format carries per-element metadata describing how each object should be virtualised in that render — but which of those instructions survive depends on the delivery path, and on the largest music service the headphone render is the platform’s own rather than the one you monitored. The practical consequence: check your mix on headphones through the actual service before you decide it is finished.
What this tool refuses to pretend
It reads the open, published parts of the file: the container defined by ITU-R BS.2088 and the metadata model defined by ITU-R BS.2076, both publicly available recommendations. It does not render. The algorithm that turns objects into speaker feeds is proprietary and licensed, so no browser can produce a conformant render, and a loudness figure taken from an unlicensed approximation would not be a valid delivery measurement. Nor does it encode or decode the broadcast codecs used to carry immersive music, which are patented. A validator that admits its boundary is more useful than a renderer that lies about one.