Training data,
on the record.
This documentation provides a high-level summary of the datasets used in the development of the generative artificial intelligence systems made available by IAH.FIT Inc. (doing business as HitZERØ), in accordance with California Civil Code Section 3111.
The datasets.Not the machine.
It describes the datasets used to train the system. It does not describe the system’s proprietary model architecture, conditioning methods, or generation processes, which are not required to be disclosed under Section 3111 and are maintained as trade secrets.
The corpus in four figures.
Distinct works in the training corpus as of this documentation date.
Created, commissioned, or acquired across this window, and continuing.
When the datasets were first used to develop the system.
No customer prompts, uploads, recordings, or outputs are used to train.
Eleven disclosures.
- 01Sources or owners of the datasets
- 02How the datasets further the intended purpose of the system
- 03Number of data points and types of data points
- 04Whether the datasets include data protected by copyright, trademark, or patent, or are entirely in the public domain
- 05Whether the datasets were purchased or licensed
- 06Whether the datasets include personal information
- 07Whether the datasets include aggregate consumer information
- 08Cleaning, processing, or other modification of the datasets
- 09Time period during which the data was collected
- 10Dates the datasets were first used in development of the system
- 11Use of synthetic data generation
Section 3111, disclosed in full.
Sources or owners of the datasets
The datasets consist of:
- Original musical works created in-house and owned by IAH.FIT Inc.;
- Musical works commissioned from third-party creators under agreements that assigned intellectual property rights to IAH.FIT Inc. upon full payment;
- Datasets obtained under paid commercial licenses that expressly authorize use for machine learning and model training;
- Openly licensed datasets used in accordance with their applicable license terms; and
- Original works generated by IAH.FIT Inc.’s own systems and owned by the company, which have been reused as training input (see Sections 3 and 11).
How the datasets further the intended purpose of the system
The datasets consist of musical audio recordings and associated descriptive metadata. They were selected and organized to enable the system to generate original instrumental and produced music responsive to specified moods, intentions, and functional states.
Number of data points and types of data points
The datasets comprise approximately 43,500 audio works and associated metadata. This figure is a high-level estimate of the documented training corpus as of the date of this documentation. It counts distinct audio works and does not separately count derivative files such as isolated stems or symbolic (MIDI) data, which are treated as associated metadata. Because the company incorporates its own model-generated works as ongoing training input (see Section 11), the corpus continues to grow over time.
Data points include full mixes, isolated stems, symbolic (MIDI) data, tempo and key information, instrumentation descriptors, and intention tags describing mood and functional state.
Complete stereo masters of each audio work.
Separated instrument and element tracks.
MIDI representations of musical material.
Measured rhythmic and harmonic values.
Descriptors of the instruments present.
Descriptors of mood and functional state.
Whether the datasets include data protected by copyright, trademark, or patent, or are entirely in the public domain
The datasets include material in which IAH.FIT Inc. holds copyright and material licensed from third-party rights holders under agreements that permit training use, including material used under Creative Commons attribution licenses. The datasets are not entirely in the public domain. To the company’s knowledge, they do not include material dedicated to the public domain. Openly licensed material used in training remains subject to the terms of the applicable licenses (including attribution requirements) and is not dedicated to the public domain.
Whether the datasets were purchased or licensed
Portions of the datasets were developed in-house and are owned by IAH.FIT Inc. Portions were commissioned by IAH.FIT Inc. under work-for-hire or assignment agreements. Portions were obtained through paid commercial licenses that expressly authorize machine-learning and training use. Portions consist of openly licensed recordings used under their applicable license terms.
Whether the datasets include personal information (Cal. Civ. Code § 1798.140(v))
The datasets contain limited personal information consisting of creator names and attribution credits associated with commissioned works, the name of the company’s founder as composer of certain in-house works, and academic-author attribution required by the open licenses governing certain academic recordings. Performing artists appearing in the openly licensed academic material are anonymized and are not identified by name.
Whether the datasets include aggregate consumer information (Cal. Civ. Code § 1798.140(b))
The datasets contain no aggregate consumer information as defined in Civil Code section 1798.140(b).
Cleaning, processing, or other modification of the datasets
The datasets underwent stem separation and structuring, structured metadata tagging (including mood, intention, tempo, key, and instrumentation descriptors), automated dataset validation and quality checks, de-duplication of redundant records, and normalization of record dates. Audio inputs were organized and standardized into consistent formats suitable for model training. These steps were performed to standardize inputs and to associate each work with structured descriptors of mood and functional state.
Stem separation and structuring of the source audio.
Structured tagging including mood, intention, tempo, key, and instrumentation descriptors.
Automated dataset validation and quality checks.
De-duplication of redundant records across the corpus.
Normalization of record dates into a consistent form.
Audio inputs organized and standardized into consistent formats suitable for model training.
Time period during which the data was collected
The data was created, commissioned, or acquired between 2018 and the present, and continues to be updated.
Dates the datasets were first used in development of the system
The datasets were first used in system development beginning in May 2025.
Use of synthetic data generation
Synthetic data generation was used in the development of these systems and continues to be used. The synthetic data consists of original musical works generated by IAH.FIT Inc.’s own systems and owned by the company. These works are reused as training input, and outputs of the system are, on an ongoing basis, reintroduced as further training input. Synthetic data is used for the purpose of expanding and refining the training corpus and improving the system’s responsiveness to mood, intention, and functional-state descriptors. Synthetic data used in training is owned by IAH.FIT Inc. and is treated as machine-generated; it is not represented as human-composed.
A living record.
This documentation is provided for transparency pursuant to California Civil Code Section 3111. It will be reviewed and updated as the systems are substantially modified.
Published by IAH.FIT Inc. (doing business as HitZERØ). Last updated July 29, 2026.