Skip to content
Cal. Civ. Code § 3111

Training data,
on the record.

This documentation provides a high-level summary of the datasets used in the development of the generative artificial intelligence systems made available by IAH.FIT Inc. (doing business as HitZERØ), in accordance with California Civil Code Section 3111.

Last updated July 29, 2026
Scope

The datasets.Not the machine.

It describes the datasets used to train the system. It does not describe the system’s proprietary model architecture, conditioning methods, or generation processes, which are not required to be disclosed under Section 3111 and are maintained as trade secrets.

At a glance

The corpus in four figures.

43,500
Documented audio works

Distinct works in the training corpus as of this documentation date.

2018–Present
Collection period

Created, commissioned, or acquired across this window, and continuing.

May 2025
First used in development

When the datasets were first used to develop the system.

Zero
Customer data in training

No customer prompts, uploads, recordings, or outputs are used to train.

The record

Section 3111, disclosed in full.

01

Sources or owners of the datasets

The datasets consist of:

  • Original musical works created in-house and owned by IAH.FIT Inc.;
  • Musical works commissioned from third-party creators under agreements that assigned intellectual property rights to IAH.FIT Inc. upon full payment;
  • Datasets obtained under paid commercial licenses that expressly authorize use for machine learning and model training;
  • Openly licensed datasets used in accordance with their applicable license terms; and
  • Original works generated by IAH.FIT Inc.’s own systems and owned by the company, which have been reused as training input (see Sections 3 and 11).
02

How the datasets further the intended purpose of the system

The datasets consist of musical audio recordings and associated descriptive metadata. They were selected and organized to enable the system to generate original instrumental and produced music responsive to specified moods, intentions, and functional states.

03

Number of data points and types of data points

The datasets comprise approximately 43,500 audio works and associated metadata. This figure is a high-level estimate of the documented training corpus as of the date of this documentation. It counts distinct audio works and does not separately count derivative files such as isolated stems or symbolic (MIDI) data, which are treated as associated metadata. Because the company incorporates its own model-generated works as ongoing training input (see Section 11), the corpus continues to grow over time.

Data points include full mixes, isolated stems, symbolic (MIDI) data, tempo and key information, instrumentation descriptors, and intention tags describing mood and functional state.

Full mixes

Complete stereo masters of each audio work.

Isolated stems

Separated instrument and element tracks.

Symbolic data

MIDI representations of musical material.

Tempo and key

Measured rhythmic and harmonic values.

Instrumentation

Descriptors of the instruments present.

Intention tags

Descriptors of mood and functional state.

04

Whether the datasets include data protected by copyright, trademark, or patent, or are entirely in the public domain

The datasets include material in which IAH.FIT Inc. holds copyright and material licensed from third-party rights holders under agreements that permit training use, including material used under Creative Commons attribution licenses. The datasets are not entirely in the public domain. To the company’s knowledge, they do not include material dedicated to the public domain. Openly licensed material used in training remains subject to the terms of the applicable licenses (including attribution requirements) and is not dedicated to the public domain.

05

Whether the datasets were purchased or licensed

Portions of the datasets were developed in-house and are owned by IAH.FIT Inc. Portions were commissioned by IAH.FIT Inc. under work-for-hire or assignment agreements. Portions were obtained through paid commercial licenses that expressly authorize machine-learning and training use. Portions consist of openly licensed recordings used under their applicable license terms.

06

Whether the datasets include personal information (Cal. Civ. Code § 1798.140(v))

The datasets contain limited personal information consisting of creator names and attribution credits associated with commissioned works, the name of the company’s founder as composer of certain in-house works, and academic-author attribution required by the open licenses governing certain academic recordings. Performing artists appearing in the openly licensed academic material are anonymized and are not identified by name.

07

Whether the datasets include aggregate consumer information (Cal. Civ. Code § 1798.140(b))

The datasets contain no aggregate consumer information as defined in Civil Code section 1798.140(b).

08

Cleaning, processing, or other modification of the datasets

The datasets underwent stem separation and structuring, structured metadata tagging (including mood, intention, tempo, key, and instrumentation descriptors), automated dataset validation and quality checks, de-duplication of redundant records, and normalization of record dates. Audio inputs were organized and standardized into consistent formats suitable for model training. These steps were performed to standardize inputs and to associate each work with structured descriptors of mood and functional state.

Stem separation

Stem separation and structuring of the source audio.

Metadata tagging

Structured tagging including mood, intention, tempo, key, and instrumentation descriptors.

Validation

Automated dataset validation and quality checks.

De-duplication

De-duplication of redundant records across the corpus.

Date normalization

Normalization of record dates into a consistent form.

Format standardization

Audio inputs organized and standardized into consistent formats suitable for model training.

09

Time period during which the data was collected

The data was created, commissioned, or acquired between 2018 and the present, and continues to be updated.

10

Dates the datasets were first used in development of the system

The datasets were first used in system development beginning in May 2025.

11

Use of synthetic data generation

Synthetic data generation was used in the development of these systems and continues to be used. The synthetic data consists of original musical works generated by IAH.FIT Inc.’s own systems and owned by the company. These works are reused as training input, and outputs of the system are, on an ongoing basis, reintroduced as further training input. Synthetic data is used for the purpose of expanding and refining the training corpus and improving the system’s responsiveness to mood, intention, and functional-state descriptors. Synthetic data used in training is owned by IAH.FIT Inc. and is treated as machine-generated; it is not represented as human-composed.

Maintained

A living record.

This documentation is provided for transparency pursuant to California Civil Code Section 3111. It will be reviewed and updated as the systems are substantially modified.

Terms of Service

Published by IAH.FIT Inc. (doing business as HitZERØ). Last updated July 29, 2026.