Mini Audition

Архітектура

Нотатки розробника англійською.

A native macOS audio editor in SwiftUI, AppKit and Metal (Swift 6, macOS 15+). It has two views over one set of tracks:

  • Waveform: one track at a time, with sample-accurate editing, effects, markers, loudness, a spectrogram and transcripts.
  • Multitrack: tracks placed as clips on lanes, mixed through buses and Master with effects, then mixed down.

About 25k lines of Swift in Sources/, plus around 385 tests in Tests/ (Swift Testing). The Xcode project is generated from project.yml with XcodeGen.

Source layout

Folder What’s in it
Sources/App @main app and AppDelegate, menus (EditorCommands, FileCommands), Now Playing / media keys, and the DEBUG screenshot and PROFILING modes
Sources/Model EditorModel (the app’s single model), AudioTrack, Viewport, markers, transcripts, preferences, effect sessions and chains
Sources/Model/Editor EditorModel+*.swift: one extension per feature (transport, editing, effects, export, recording, documents, multitrack, …)
Sources/Model/Multitrack Timeline (lanes, clips, buses, markers, zoom, undo), TimelineMixer (the render-thread mix), TimelinePlayback, LevelMeters
Sources/Model/Session The autosaved session, the .miniaudition bundle format, the timeline’s saved form
Sources/Audio Decoded audio (AudioData), loading, editing, playback engine, recording, export, loudness, silence detection, speech-to-text
Sources/Audio/Effects Effect DSP: one …Processor per effect, EffectKind (the catalogue), parameters, presets, the spectrum analysis
Sources/Audio/Plugins Third-party effects: Audio Units (AudioUnitPlugin, AudioUnitProcessor)
Sources/Renderer Metal: the shared context and pipelines, the waveform/ruler/meter/marker/HUD renderers, glyph atlas, the GPU spectrogram, the spectrum analyzer, .metal shaders
Sources/Views SwiftUI and AppKit views: window layout, sidebar, transport bar, inspector, settings
Sources/Views/Canvas EditorCanvasView, the waveform view’s single MTKView (drawing, mouse, keyboard, context menu)
Sources/Views/Multitrack The multitrack view: lanes, ruler, overview, mixer, panel, plus the layer-driven playhead, clock and meters
Sources/Views/Effects The effect chain strip, knobs, EQ editor, reduction meter, spectrum display
Sources/SessionPreview + QuickLook/ The Quick Look extension that previews .miniaudition files in Finder (sandboxed)
bin/ Build, test helpers, release (notarize, install), screenshots, profile

The big picture

flowchart TD
    UI[SwiftUI views and menus] -->|read, observed| M[EditorModel]
    UI -->|actions| M
    M --> T[AudioTrack: AudioData, markers, undo]
    M --> TL[Timeline: lanes, clips, buses, effects, undo]
    M --> PE[PlaybackEngine]
    PE -->|render thread pulls| PC{PlaybackContent}
    PC --> RS[RenderSource: a track]
    PC --> TP[TimelinePlayback: the mix]
    TP --> MX[TimelineMixer]
    MX --> FX[Effect inserts and LevelMeters]
    RS --> FX2[Effect chain preview]
    T --> GPU[Metal: waveform canvas, spectrogram]
    TL --> MV[Multitrack views and FrameDriver layers]
    M --> S[Session autosave and .miniaudition bundles]

Model

  • EditorModel is @MainActor @Observable, owned by AppDelegate so that files opened from Finder reach it before any window exists.
    • Its stored properties live in EditorModel.swift; each feature adds an extension in Model/Editor/.
    • Views read it directly. Observation re-renders only what a view actually read.
  • AudioTrack is one file: its decoded audio, viewport, selection, cursor, gain, markers, transcript and the cached spectrogram and loudness. Each track has its own UndoManager.
  • Timeline is the multitrack arrangement. It has its own UndoManager:
    • registerUndo snapshots the lanes, buses, Master gain and markers as values;
    • Edit › Undo uses that history while the multitrack view is showing.
  • Edit-time vs view state: zoom, selection and cursor aren’t undone. Effect settings aren’t part of either undo history (as in most DAWs).
  • editorMode (.waveform / .multitrack) decides what Play, Undo, Zoom and the Effects menu act on. activeTrack is nil in the multitrack view, so track commands can’t touch a track that isn’t showing.

Audio data

  • AudioData is immutable decoded audio.
    • Each channel’s samples sit in a shared-storage MTLBuffer, and that single copy feeds both the GPU (drawing) and the render thread (playback).
    • With it comes a min/max + energy peak pyramid, so drawing any zoom level reads a few thousand entries, not millions of samples.
  • Edits (AudioTrack.replace) build a new AudioData, and the old one stays valid: it may still be on screen, playing, or held by undo.
    • Undo keeps the removed span as an AudioClip, the same type the clipboard uses.
    • Markers shift with the audio.
  • Loading (AudioLoader) decodes with AVAudioFile in chunks, with progress, off the main thread.
  • Recording (AudioRecorder) runs its own AVAudioEngine, so the playback engine is never disturbed.
    • It streams to a file (RecordingWriter).
    • LiveAudio extends a growing AudioData and its peaks incrementally, so the take draws live at constant cost.
  • Export (AudioExporter, ExportFormat): gain, then resampling if asked (AVAudioConverter, mastering quality), TPDF dither for 16 and 24 bits, the clamp to full scale (not for float), and AVAudioFile. Broadcast WAV’s bext chunk is written into the FLLR filler Core Audio leaves before the audio (no copying), or appended.
  • Tags (AudioTags): written into the finished file, each format its own way: an ID3v2.3 tag prepended to MP3 (the encoders are told to write none) and appended as a chunk to WAV (with a RIFF INFO list) and AIFF; FLAC’s Vorbis comment and picture blocks replace the encoder’s, the frames copied after them; M4A goes through a passthrough AVAssetExportSession with iTunes metadata (no re-encoding).
  • MP3 (MP3Encoder): macOS decodes MP3 but can’t encode it, and LAME isn’t bundled (LGPL). An installed lame or LAME-enabled ffmpeg (Homebrew, MacPorts paths: an app opened from Finder has no shell PATH) encodes a 24-bit WAV the export writes first, as its own process, so the hardened runtime’s library validation doesn’t apply. Core Audio’s encoder is used instead if a macOS version gains one.
  • Loudness (Loudness, TruePeak, LoudnessReport): BS.1770 from the K-weighted power of 100 ms segments, which also give the momentary (400 ms) and short-term (3 s) loudness and the loudness range (EBU Tech 3342). The true peak oversamples 4× with a polyphase windowed sinc run by vDSP_conv: about half a second per stereo hour. Gain is linear, so a report is measured once and shifted by the track’s gain.
  • Batch (BatchProcessing, BatchJob, EditorModel+Batch): one file at a time, decoded, run through fresh copies of the effect panel’s chain (padded for the tails), measured with LoudnessReport, leveled (loudness, held under a true-peak ceiling, or the peak to it) and exported, without becoming tracks.
  • Memory: AudioMemory totals the decoded buffers for the status bar. Recording snapshots share buffers and are counted once.

Playback

  • PlaybackEngine wraps AVAudioEngine with an AVAudioSourceNode that pulls straight from memory, followed by an AVAudioUnitTimePitch for playback speed with the pitch kept.
  • What it plays is either of two PlaybackContents:
    • RenderSource: a track’s AudioData, with gain ramps, ranges and looping, plus the effect chain preview through an EffectInsert.
    • TimelinePlayback: calls TimelineMixer.render for the timeline.
  • Render-thread rules (stated at the top of each type):
    • no allocation, no locks the main thread can hold, no Objective-C, no I/O;
    • state crosses threads only through Synchronization atomics;
    • snapshots and effect processors are swapped with Mutex.withLockIfAvailable. If the UI holds the lock at that moment, one buffer plays dry or silent; the audio thread never waits.
  • Latency: the render position runs ahead of the speakers by the output latency, plus the time-stretch latency at other speeds. The engine reads that latency once per play (asking the HAL is slow) and exposes:
    • displayPosition, for the playhead;
    • latency, for meters that must show what’s heard. That’s tens of ms on speakers and about 250 ms on AirPods.

Multitrack

  • The Timeline holds:
    • lanes of clips, each a stretch of an AudioData with fades;
    • buses and Master, each with an EffectChain;
    • mix markers.
  • Every change rebuilds a value TimelineMixer.Snapshot. TimelinePlayback swaps it in under its mutex, and the old snapshot is released on the main thread, so a removed clip’s audio is never freed on the render thread.
  • TimelineMixer mixes block by block:
    1. lanes go into buses or straight into Master;
    2. each bus runs its effect insert and gain;
    3. Master runs its effects and gain.

    Delay compensation: an effect can play late (the Clean pitch shifter, by its slice length; EffectProcessor.latency). Each bus reports its chain’s latency as it runs, and TimelineMixer.Compensation delays the lanes going straight to Master and every other bus to the latest one (as of the block before), so they reach Master together. Timeline.latency (the latest bus plus Master’s effects) is added to the playhead’s latency; meters stamp late paths back to their place on the timeline.

    Sends: a lane’s copies into other buses (TimelineSend, post-fader), added into each bus’s sum by the same add with the send’s gain; the lane’s meter shows its own output only. A bus fed by sends alone runs and counts for delay compensation like any other.

    Stretching a clip (AudioData.stretched, Timeline+Stretch): the clip’s stretch of audio through AVAudioUnitTimePitch (overlap 32) in an offline engine, rendered in the background to exactly the new length, as audio of the clip’s own. It goes in only if the clip hasn’t changed meanwhile.

    Automation (TimelineEnvelope): points in timeline frames joined by straight lines, held at the ends. A lane’s volume line is the level in dB (silence at the bottom) and replaces its fader; its balance replaces the knob. Sessions before format 3 saved volume as an offset on the fader; TimelineState.withAbsoluteVolume adds the fader in on load. A lane with automation is mixed by addAutomated, which fills a block’s gains and balance into scratch buffers first (no allocation on the render thread); lanes without any keep the faster constant-gain path.

    Mixer automation (Timeline+MixerAutomation, AutomationScale): bus and Master volume envelopes, and each effect’s automation (by parameter index, or Audio Unit address). ChainProcessor.Stage carries them; the mixer passes the block’s timeline frame to EffectInsert.process, and the chain sets each parameter for the block: a built-in one in its ParameterBank, an Audio Unit’s through scheduleParameterBlock. One editor view (EnvelopeEditor) draws every line, given a scale and closures.

    Automated controls show their lines’ values without touching the settings: AutomationClock holds the playhead (set by the 30 Hz playback timer; the cursor when stopped), and FollowingLine, ParameterKnob and the EQ read it only when there’s a line, so only those redraw as it moves. A control with a line can’t be moved meanwhile: its setting isn’t heard.

    Mute is applied when the snapshot is built. The first clip placed sets the timeline’s sample rate; later clips are resampled when placed.

  • Metering: LevelMeters keeps a ring per lane, bus and Master of peaks stamped with a monotonic “frames mixed” count. The UI takes only entries older than the output latency, so the meters move with the sound, and loops and seeks don’t matter.
  • Mixdown renders the same snapshot offline:
    • with fresh effect processors and a frozen copy of the parameters;
    • with room left for the effects’ tails;
    • rendering latency frames more and dropping the start, so it lines up with the clips;
    • into a 32-bit float WAV, added as a new track.
  • Waveform → multitrack: the waveform context menu sends the selection, regions or the whole track as clips that share the track’s audio. They’re saved as references to the track.

Effects

  • EffectKind is the catalogue: title, icon, summary, menu category, parameters, presets, processor factory and tail length. Adding an effect means:
    1. a processor;
    2. a case in EffectKind;
    3. presets;
    4. tests (the shared tests cover stability at the extremes, preview = Apply, and preset validity).
  • EffectProcessor is the DSP for one sample rate and channel count. It processes in place on the render thread and reads its knobs from a ParameterBank of atomics.
  • Readouts: the compressor, limiter and gate report gain reduction back through the bank as a short history stamped by frames processed, so the meter can show the value for what’s being heard.
  • EffectSession is one effect in a chain as the UI sees it: values, bypass, and a cached preview processor so echoes ring on when the chain is reordered. EffectChain orders sessions. A ChainProcessor runs them in sequence, skipping bypassed ones.
  • Preview vs Apply: the preview runs the live processors on the live bank. Apply and mixdown build fresh processors with a snapshot of the values, so turning a knob mid-render changes nothing.
  • Visual effects (the Spectrum Analyzer) only tap the audio into a SampleTap. Apply and mixdown skip them.
  • Noise Reduction (SpectralDenoiser, NoiseProfile): spectral subtraction in 43 ms Hann slices overlapping 4×, against a captured noise print (mean power per bin, read off by frequency at other rates and sizes) or a minimum-statistics estimate of a smoothed power. The gain is decided on power averaged over neighbouring bins and recent slices, which keeps leftover noise from twinkling. The print is session data, not a knob: EffectSession.makeProcessor passes it, and the mixer saves it with the effect.
  • Spectral editing (SpectralEditing, SpectralSelection): a band (Hz) kept on AudioTrack with the time selection it was drawn with, valid only while that’s still the selection. The edit runs the stretch plus a slice either side through 43 ms Hann slices, masks bins by how much of each slice’s window is inside the selection and how far each bin is from the band (soft edges, no clicks), and overlap-adds with the window energy divided out, so outside the mask the audio comes back as it was. Filling scales each bin toward the average magnitude just before and after (interpolated), never up.
  • Audio Units are one EffectKind (.audioUnit) whose session holds the plugin (AudioUnitInfo, its instance, and its fullState as a binary plist) instead of values.
    • Instances load out of process (.loadOutOfProcess, AUv2 ones too), so a crashing plugin doesn’t take the app down, and need no library-validation entitlement.
    • Loading is async: the session is in the chain at once (isLoading, passed through), and Apply and mixdown wait until it’s done. A restored one that won’t load stays as a placeholder keeping its saved state.
    • AudioUnitProcessor wraps the render block: both buses enabled (bridged AUv2 units fail with -10876 otherwise), blocks of at most 4096 frames, input pulled from a copy. The preview’s processors share the unit, so AudioUnitRenderGate lets only the newest one render it; the others pass through, never touching a unit being reconfigured.
    • Apply and mixdown render a fresh instance with the preview’s fullState copied over.
    • Settings are captured (debounced) from parameter changes and again before saving; equal state keeps its old bytes, so loading a session doesn’t mark it edited.
    • The plugin’s own view opens in a floating window (PluginWindows), or generic sliders for plugins without one.

Stem separation

  • A separate program. Extract Stems runs Meta’s Demucs (htdemucs) on MLX, which is most of the work’s size and has no Intel build, so it isn’t in the app: the helper Mini Audition Stems (Stems/main.swift, built from the same project.yml) is downloaded from the website the first time it’s needed (StemTool).
    • The app runs it with a model folder, the input and an output folder; it prints progress lines and done, and cancelling terminates it.
    • Its version (STEMS_VERSION) is in both apps’ Info.plist; the app takes the zip that matches its own. A downloaded helper only runs if it’s signed by the app’s developer.
    • It’s always built optimized: unoptimized MLX code is far too slow.
  • The model (StemModel) comes from Hugging Face once, pinned to one revision, every file checked against its size and SHA-256 before it’s moved into place in one step.
  • Downloads report progress by sampling countOfBytesReceived: neither the async URLSession.download nor a task’s Progress does.

Drawing

The rule, learned the hard way: nothing in SwiftUI changes per frame while playing. Every animated element is either Metal or Core Animation layers. Measure changes with bin/profile (below).

  • Waveform view: EditorCanvasView is one MTKView. Ruler, waveform or spectrogram, markers, overview, level meter and the time readout are regions of the same drawable (EditorLayout), drawn from one playhead value per frame.
    • Renderers: WaveformRenderer, RulerRenderer, MarkerRenderer, MeterRenderer, HUDRenderer.
    • Text: GlyphAtlas provides it.
    • Spectrogram: a GPU STFT compute pass (Spectrogram), stored as one byte per bin.
    • Shared drawing: QuadDrawing gives every renderer instanced rectangles, glyphs and coloured triangles. The pipelines live once in MetalContext.
  • Multitrack view: this is SwiftUI (lanes, mixer, panels), except for what moves while playing. A single FrameDriver (a display link capped at 60 Hz, running only while the timeline plays) updates layer-backed views in one transaction:
    • TimelinePositionLines (playhead and cursor);
    • TimelineClock;
    • LevelMeter;
    • LevelLED.

    This took multitrack playback from about 27% to about 5% of a core.

  • Effect panels: the knobs and the EQ editor are SwiftUI. The Spectrum Analyzer is an MTKView (SpectrumRenderer, geometry built on the CPU by MeshBuilder) that pauses once its bars have fallen.
  • Themes: CanvasTheme holds a few base colours and derives the rest. The rulers always follow the system’s light or dark look.

Threads

Thread Does
Main Model, UI, Metal encoding, the display links, meters reading their stamped history
Audio render (Core Audio IO thread) RenderSource / TimelineMixer, effect processors, meter and tap writes
Background tasks (Task.detached) Decoding, loudness, silence detection, Apply, mixdown, export, session writes, transcription
GPU Drawing, the spectrogram STFT
Recording engine Input tap, writing the file

Swift 6 strict concurrency is on. Render-thread types are Sendable through atomics, or @unchecked Sendable with the reasoning documented where they are.

Persistence

  • Autosave: the untitled session (SessionState: tracks by path plus bookmark, their settings, the timeline) goes to Application Support. Saves are debounced, triggered by observation, and restored at launch. Hold ⇧ or pass --no-session to start blank.
  • Saved sessions are .miniaudition packages (SessionBundle; format 1 without a timeline, 2 with one, 3 with volume automation, so older versions refuse what they’d get wrong):
    • Session.json with the settings;
    • a copy of every track in Audio/, with edited tracks as 32-bit float WAV (files the app made, such as recordings, are moved in rather than copied);
    • clips’ own audio in Audio/Clips/;
    • a snapshot at every save in Snapshots/<date>/: its Session.json and a JPEG of the window (WindowSnapshot; Metal views render a frame of their own for it). Their paths point at the bundle’s audio, so a save carries over the files older snapshots still need, hard-linked so that unchanged audio is stored once.

    Older app versions refuse newer formats instead of dropping data. New fields are optional, so older sessions still open.

  • Preferences live in UserDefaults (Preferences): last-used effect settings, themes, recording and transcription options.
  • Quick Look: the SessionPreview extension reads a bundle’s Session.json to preview it in Finder, with the latest snapshot’s picture of the window at the top.

System integration

  • Media controls: NowPlaying handles AirPods, media keys and Control Center. Play and pause work as Space does, next and previous jump between markers, and seeking works from Now Playing.
  • Space anywhere: SpaceBarCatcher is a window-wide key monitor, so Space plays even when nothing in the editor has focus.
  • Drag out: ⌥-drag a selection to Finder; AudioFilePromise writes the file after the drop.
  • On-device AI: speech-to-text uses SpeechAnalyzer (macOS 26). Titles, summaries and sections use Foundation Models (Apple Intelligence).

Tooling

Command Does
bin/build [--run] Debug build into build/Debug
xcodegen generate Regenerate the project. Needed after adding files: project.yml globs folders.
xcodebuild test -project MiniAudition.xcodeproj -scheme MiniAudition -destination 'platform=macOS' The test suite (it runs inside the app as its host)
bin/screenshots DEBUG screenshot mode: stages scenes and captures the README images
bin/profile <session> [--without part] [--analyzer] [--idle] Optimized build with PROFILING. Plays a session, attaches Time Profiler, prints a per-thread / per-area breakdown (bin/profile-report). --without stops one animated part to measure its cost by its absence.
bin/notarize, bin/install Release DMG (signed, notarized, stapled) and installing it (see AGENTS.md)

Conventions and pitfalls

  • One feature per EditorModel+*.swift. Keep files under about 300 lines and views small. Comments explain why, not what.
  • Layout:
    • measured sizes never drive layout;
    • no ViewThatFits;
    • the inspector is a custom column, not .inspector;
    • width-dependent layouts are picked from hidden copies at their ideal sizes.
  • Layer-backed NSViews set isFlipped. AppKit overrides a layer’s own isGeometryFlipped.
  • SF Symbols: check a name with NSImage(systemSymbolName:) before using it. On macOS 27, AppKit menu items need preferredImageVisibility = .visible to show their icons.
  • Undo in tests: set groupsByEvent = false and wrap registering calls in undoManager.step { }. Ending a group by hand hangs the test host.