A native macOS audio editor in SwiftUI, AppKit and Metal (Swift 6, macOS 15+). It has two views over one set of tracks:
- Waveform: one track at a time, with sample-accurate editing, effects, markers, loudness, a spectrogram and transcripts.
- Multitrack: tracks placed as clips on lanes, mixed through buses and Master with effects, then mixed down.
About 25k lines of Swift in Sources/, plus around 385 tests in Tests/ (Swift Testing).
The Xcode project is generated from project.yml with XcodeGen.
Source layout
| Folder | What’s in it |
|---|---|
Sources/App |
@main app and AppDelegate, menus (EditorCommands, FileCommands), Now Playing / media keys, and the DEBUG screenshot and PROFILING modes |
Sources/Model |
EditorModel (the app’s single model), AudioTrack, Viewport, markers, transcripts, preferences, effect sessions and chains |
Sources/Model/Editor |
EditorModel+*.swift: one extension per feature (transport, editing, effects, export, recording, documents, multitrack, …) |
Sources/Model/Multitrack |
Timeline (lanes, clips, buses, markers, zoom, undo), TimelineMixer (the render-thread mix), TimelinePlayback, LevelMeters |
Sources/Model/Session |
The autosaved session, the .miniaudition bundle format, the timeline’s saved form |
Sources/Audio |
Decoded audio (AudioData), loading, editing, playback engine, recording, export, loudness, silence detection, speech-to-text |
Sources/Audio/Effects |
Effect DSP: one …Processor per effect, EffectKind (the catalogue), parameters, presets, the spectrum analysis |
Sources/Audio/Plugins |
Third-party effects: Audio Units (AudioUnitPlugin, AudioUnitProcessor) |
Sources/Renderer |
Metal: the shared context and pipelines, the waveform/ruler/meter/marker/HUD renderers, glyph atlas, the GPU spectrogram, the spectrum analyzer, .metal shaders |
Sources/Views |
SwiftUI and AppKit views: window layout, sidebar, transport bar, inspector, settings |
Sources/Views/Canvas |
EditorCanvasView, the waveform view’s single MTKView (drawing, mouse, keyboard, context menu) |
Sources/Views/Multitrack |
The multitrack view: lanes, ruler, overview, mixer, panel, plus the layer-driven playhead, clock and meters |
Sources/Views/Effects |
The effect chain strip, knobs, EQ editor, reduction meter, spectrum display |
Sources/SessionPreview + QuickLook/ |
The Quick Look extension that previews .miniaudition files in Finder (sandboxed) |
bin/ |
Build, test helpers, release (notarize, install), screenshots, profile |
The big picture
flowchart TD
UI[SwiftUI views and menus] -->|read, observed| M[EditorModel]
UI -->|actions| M
M --> T[AudioTrack: AudioData, markers, undo]
M --> TL[Timeline: lanes, clips, buses, effects, undo]
M --> PE[PlaybackEngine]
PE -->|render thread pulls| PC{PlaybackContent}
PC --> RS[RenderSource: a track]
PC --> TP[TimelinePlayback: the mix]
TP --> MX[TimelineMixer]
MX --> FX[Effect inserts and LevelMeters]
RS --> FX2[Effect chain preview]
T --> GPU[Metal: waveform canvas, spectrogram]
TL --> MV[Multitrack views and FrameDriver layers]
M --> S[Session autosave and .miniaudition bundles]
Model
EditorModelis@MainActor @Observable, owned byAppDelegateso that files opened from Finder reach it before any window exists.- Its stored properties live in
EditorModel.swift; each feature adds an extension inModel/Editor/. - Views read it directly. Observation re-renders only what a view actually read.
- Its stored properties live in
AudioTrackis one file: its decoded audio, viewport, selection, cursor, gain, markers, transcript and the cached spectrogram and loudness. Each track has its ownUndoManager.Timelineis the multitrack arrangement. It has its ownUndoManager:registerUndosnapshots the lanes, buses, Master gain and markers as values;- Edit › Undo uses that history while the multitrack view is showing.
- Edit-time vs view state: zoom, selection and cursor aren’t undone. Effect settings aren’t part of either undo history (as in most DAWs).
editorMode(.waveform/.multitrack) decides what Play, Undo, Zoom and the Effects menu act on.activeTrackis nil in the multitrack view, so track commands can’t touch a track that isn’t showing.
Audio data
AudioDatais immutable decoded audio.- Each channel’s samples sit in a shared-storage
MTLBuffer, and that single copy feeds both the GPU (drawing) and the render thread (playback). - With it comes a min/max + energy peak pyramid, so drawing any zoom level reads a few thousand entries, not millions of samples.
- Each channel’s samples sit in a shared-storage
- Edits (
AudioTrack.replace) build a newAudioData, and the old one stays valid: it may still be on screen, playing, or held by undo.- Undo keeps the removed span as an
AudioClip, the same type the clipboard uses. - Markers shift with the audio.
- Undo keeps the removed span as an
- Loading (
AudioLoader) decodes withAVAudioFilein chunks, with progress, off the main thread. - Recording (
AudioRecorder) runs its ownAVAudioEngine, so the playback engine is never disturbed.- It streams to a file (
RecordingWriter). LiveAudioextends a growingAudioDataand its peaks incrementally, so the take draws live at constant cost.
- It streams to a file (
- Export (
AudioExporter,ExportFormat): gain, then resampling if asked (AVAudioConverter, mastering quality), TPDF dither for 16 and 24 bits, the clamp to full scale (not for float), andAVAudioFile. Broadcast WAV’sbextchunk is written into theFLLRfiller Core Audio leaves before the audio (no copying), or appended. - Tags (
AudioTags): written into the finished file, each format its own way: an ID3v2.3 tag prepended to MP3 (the encoders are told to write none) and appended as a chunk to WAV (with a RIFF INFO list) and AIFF; FLAC’s Vorbis comment and picture blocks replace the encoder’s, the frames copied after them; M4A goes through a passthroughAVAssetExportSessionwith iTunes metadata (no re-encoding). - MP3 (
MP3Encoder): macOS decodes MP3 but can’t encode it, and LAME isn’t bundled (LGPL). An installedlameor LAME-enabledffmpeg(Homebrew, MacPorts paths: an app opened from Finder has no shell PATH) encodes a 24-bit WAV the export writes first, as its own process, so the hardened runtime’s library validation doesn’t apply. Core Audio’s encoder is used instead if a macOS version gains one. - Loudness (
Loudness,TruePeak,LoudnessReport): BS.1770 from the K-weighted power of 100 ms segments, which also give the momentary (400 ms) and short-term (3 s) loudness and the loudness range (EBU Tech 3342). The true peak oversamples 4× with a polyphase windowed sinc run byvDSP_conv: about half a second per stereo hour. Gain is linear, so a report is measured once and shifted by the track’s gain. - Batch (
BatchProcessing,BatchJob,EditorModel+Batch): one file at a time, decoded, run through fresh copies of the effect panel’s chain (padded for the tails), measured withLoudnessReport, leveled (loudness, held under a true-peak ceiling, or the peak to it) and exported, without becoming tracks. - Memory:
AudioMemorytotals the decoded buffers for the status bar. Recording snapshots share buffers and are counted once.
Playback
PlaybackEnginewrapsAVAudioEnginewith anAVAudioSourceNodethat pulls straight from memory, followed by anAVAudioUnitTimePitchfor playback speed with the pitch kept.- What it plays is either of two
PlaybackContents:RenderSource: a track’sAudioData, with gain ramps, ranges and looping, plus the effect chain preview through anEffectInsert.TimelinePlayback: callsTimelineMixer.renderfor the timeline.
- Render-thread rules (stated at the top of each type):
- no allocation, no locks the main thread can hold, no Objective-C, no I/O;
- state crosses threads only through
Synchronizationatomics; - snapshots and effect processors are swapped with
Mutex.withLockIfAvailable. If the UI holds the lock at that moment, one buffer plays dry or silent; the audio thread never waits.
- Latency: the render position runs ahead of the speakers by the output latency, plus
the time-stretch latency at other speeds. The engine reads that latency once per play
(asking the HAL is slow) and exposes:
displayPosition, for the playhead;latency, for meters that must show what’s heard. That’s tens of ms on speakers and about 250 ms on AirPods.
Multitrack
- The
Timelineholds:- lanes of clips, each a stretch of an
AudioDatawith fades; - buses and Master, each with an
EffectChain; - mix markers.
- lanes of clips, each a stretch of an
- Every change rebuilds a value
TimelineMixer.Snapshot.TimelinePlaybackswaps it in under its mutex, and the old snapshot is released on the main thread, so a removed clip’s audio is never freed on the render thread. TimelineMixermixes block by block:- lanes go into buses or straight into Master;
- each bus runs its effect insert and gain;
- Master runs its effects and gain.
Delay compensation: an effect can play late (the Clean pitch shifter, by its slice length;
EffectProcessor.latency). Each bus reports its chain’s latency as it runs, andTimelineMixer.Compensationdelays the lanes going straight to Master and every other bus to the latest one (as of the block before), so they reach Master together.Timeline.latency(the latest bus plus Master’s effects) is added to the playhead’s latency; meters stamp late paths back to their place on the timeline.Sends: a lane’s copies into other buses (
TimelineSend, post-fader), added into each bus’s sum by the sameaddwith the send’s gain; the lane’s meter shows its own output only. A bus fed by sends alone runs and counts for delay compensation like any other.Stretching a clip (
AudioData.stretched,Timeline+Stretch): the clip’s stretch of audio throughAVAudioUnitTimePitch(overlap 32) in an offline engine, rendered in the background to exactly the new length, as audio of the clip’s own. It goes in only if the clip hasn’t changed meanwhile.Automation (
TimelineEnvelope): points in timeline frames joined by straight lines, held at the ends. A lane’s volume line is the level in dB (silence at the bottom) and replaces its fader; its balance replaces the knob. Sessions before format 3 saved volume as an offset on the fader;TimelineState.withAbsoluteVolumeadds the fader in on load. A lane with automation is mixed byaddAutomated, which fills a block’s gains and balance into scratch buffers first (no allocation on the render thread); lanes without any keep the faster constant-gain path.Mixer automation (
Timeline+MixerAutomation,AutomationScale): bus and Master volume envelopes, and each effect’sautomation(by parameter index, or Audio Unit address).ChainProcessor.Stagecarries them; the mixer passes the block’s timeline frame toEffectInsert.process, and the chain sets each parameter for the block: a built-in one in itsParameterBank, an Audio Unit’s throughscheduleParameterBlock. One editor view (EnvelopeEditor) draws every line, given a scale and closures.Automated controls show their lines’ values without touching the settings:
AutomationClockholds the playhead (set by the 30 Hz playback timer; the cursor when stopped), andFollowingLine,ParameterKnoband the EQ read it only when there’s a line, so only those redraw as it moves. A control with a line can’t be moved meanwhile: its setting isn’t heard.Mute is applied when the snapshot is built. The first clip placed sets the timeline’s sample rate; later clips are resampled when placed.
- Metering:
LevelMeterskeeps a ring per lane, bus and Master of peaks stamped with a monotonic “frames mixed” count. The UI takes only entries older than the output latency, so the meters move with the sound, and loops and seeks don’t matter. - Mixdown renders the same snapshot offline:
- with fresh effect processors and a frozen copy of the parameters;
- with room left for the effects’ tails;
- rendering
latencyframes more and dropping the start, so it lines up with the clips; - into a 32-bit float WAV, added as a new track.
- Waveform → multitrack: the waveform context menu sends the selection, regions or the whole track as clips that share the track’s audio. They’re saved as references to the track.
Effects
EffectKindis the catalogue: title, icon, summary, menu category, parameters, presets, processor factory and tail length. Adding an effect means:- a processor;
- a case in
EffectKind; - presets;
- tests (the shared tests cover stability at the extremes, preview = Apply, and preset validity).
EffectProcessoris the DSP for one sample rate and channel count. It processes in place on the render thread and reads its knobs from aParameterBankof atomics.- Readouts: the compressor, limiter and gate report gain reduction back through the bank as a short history stamped by frames processed, so the meter can show the value for what’s being heard.
EffectSessionis one effect in a chain as the UI sees it: values, bypass, and a cached preview processor so echoes ring on when the chain is reordered.EffectChainorders sessions. AChainProcessorruns them in sequence, skipping bypassed ones.- Preview vs Apply: the preview runs the live processors on the live bank. Apply and mixdown build fresh processors with a snapshot of the values, so turning a knob mid-render changes nothing.
- Visual effects (the Spectrum Analyzer) only tap the audio into a
SampleTap. Apply and mixdown skip them. - Noise Reduction (
SpectralDenoiser,NoiseProfile): spectral subtraction in 43 ms Hann slices overlapping 4×, against a captured noise print (mean power per bin, read off by frequency at other rates and sizes) or a minimum-statistics estimate of a smoothed power. The gain is decided on power averaged over neighbouring bins and recent slices, which keeps leftover noise from twinkling. The print is session data, not a knob:EffectSession.makeProcessorpasses it, and the mixer saves it with the effect. - Spectral editing (
SpectralEditing,SpectralSelection): a band (Hz) kept onAudioTrackwith the time selection it was drawn with, valid only while that’s still the selection. The edit runs the stretch plus a slice either side through 43 ms Hann slices, masks bins by how much of each slice’s window is inside the selection and how far each bin is from the band (soft edges, no clicks), and overlap-adds with the window energy divided out, so outside the mask the audio comes back as it was. Filling scales each bin toward the average magnitude just before and after (interpolated), never up. - Audio Units are one
EffectKind(.audioUnit) whose session holds the plugin (AudioUnitInfo, its instance, and itsfullStateas a binary plist) instead of values.- Instances load out of process (
.loadOutOfProcess, AUv2 ones too), so a crashing plugin doesn’t take the app down, and need no library-validation entitlement. - Loading is async: the session is in the chain at once (
isLoading, passed through), and Apply and mixdown wait until it’s done. A restored one that won’t load stays as a placeholder keeping its saved state. AudioUnitProcessorwraps the render block: both buses enabled (bridged AUv2 units fail with -10876 otherwise), blocks of at most 4096 frames, input pulled from a copy. The preview’s processors share the unit, soAudioUnitRenderGatelets only the newest one render it; the others pass through, never touching a unit being reconfigured.- Apply and mixdown render a fresh instance with the preview’s
fullStatecopied over. - Settings are captured (debounced) from parameter changes and again before saving; equal state keeps its old bytes, so loading a session doesn’t mark it edited.
- The plugin’s own view opens in a floating window (
PluginWindows), or generic sliders for plugins without one.
- Instances load out of process (
Stem separation
- A separate program. Extract Stems runs Meta’s Demucs (
htdemucs) on MLX, which is most of the work’s size and has no Intel build, so it isn’t in the app: the helperMini Audition Stems(Stems/main.swift, built from the sameproject.yml) is downloaded from the website the first time it’s needed (StemTool).- The app runs it with a model folder, the input and an output folder; it prints
progresslines anddone, and cancelling terminates it. - Its version (
STEMS_VERSION) is in both apps’Info.plist; the app takes the zip that matches its own. A downloaded helper only runs if it’s signed by the app’s developer. - It’s always built optimized: unoptimized MLX code is far too slow.
- The app runs it with a model folder, the input and an output folder; it prints
- The model (
StemModel) comes from Hugging Face once, pinned to one revision, every file checked against its size and SHA-256 before it’s moved into place in one step. - Downloads report progress by sampling
countOfBytesReceived: neither the asyncURLSession.downloadnor a task’sProgressdoes.
Drawing
The rule, learned the hard way: nothing in SwiftUI changes per frame while playing.
Every animated element is either Metal or Core Animation layers. Measure changes with
bin/profile (below).
- Waveform view:
EditorCanvasViewis oneMTKView. Ruler, waveform or spectrogram, markers, overview, level meter and the time readout are regions of the same drawable (EditorLayout), drawn from one playhead value per frame.- Renderers:
WaveformRenderer,RulerRenderer,MarkerRenderer,MeterRenderer,HUDRenderer. - Text:
GlyphAtlasprovides it. - Spectrogram: a GPU STFT compute pass (
Spectrogram), stored as one byte per bin. - Shared drawing:
QuadDrawinggives every renderer instanced rectangles, glyphs and coloured triangles. The pipelines live once inMetalContext.
- Renderers:
- Multitrack view: this is SwiftUI (lanes, mixer, panels), except for what moves while
playing. A single
FrameDriver(a display link capped at 60 Hz, running only while the timeline plays) updates layer-backed views in one transaction:TimelinePositionLines(playhead and cursor);TimelineClock;LevelMeter;LevelLED.
This took multitrack playback from about 27% to about 5% of a core.
- Effect panels: the knobs and the EQ editor are SwiftUI. The Spectrum Analyzer is an
MTKView(SpectrumRenderer, geometry built on the CPU byMeshBuilder) that pauses once its bars have fallen. - Themes:
CanvasThemeholds a few base colours and derives the rest. The rulers always follow the system’s light or dark look.
Threads
| Thread | Does |
|---|---|
| Main | Model, UI, Metal encoding, the display links, meters reading their stamped history |
| Audio render (Core Audio IO thread) | RenderSource / TimelineMixer, effect processors, meter and tap writes |
Background tasks (Task.detached) |
Decoding, loudness, silence detection, Apply, mixdown, export, session writes, transcription |
| GPU | Drawing, the spectrogram STFT |
| Recording engine | Input tap, writing the file |
Swift 6 strict concurrency is on. Render-thread types are Sendable through atomics, or
@unchecked Sendable with the reasoning documented where they are.
Persistence
- Autosave: the untitled session (
SessionState: tracks by path plus bookmark, their settings, the timeline) goes to Application Support. Saves are debounced, triggered by observation, and restored at launch. Hold ⇧ or pass--no-sessionto start blank. - Saved sessions are
.miniauditionpackages (SessionBundle; format 1 without a timeline, 2 with one, 3 with volume automation, so older versions refuse what they’d get wrong):Session.jsonwith the settings;- a copy of every track in
Audio/, with edited tracks as 32-bit float WAV (files the app made, such as recordings, are moved in rather than copied); - clips’ own audio in
Audio/Clips/; - a snapshot at every save in
Snapshots/<date>/: itsSession.jsonand a JPEG of the window (WindowSnapshot; Metal views render a frame of their own for it). Their paths point at the bundle’s audio, so a save carries over the files older snapshots still need, hard-linked so that unchanged audio is stored once.
Older app versions refuse newer formats instead of dropping data. New fields are optional, so older sessions still open.
- Preferences live in
UserDefaults(Preferences): last-used effect settings, themes, recording and transcription options. - Quick Look: the
SessionPreviewextension reads a bundle’sSession.jsonto preview it in Finder, with the latest snapshot’s picture of the window at the top.
System integration
- Media controls:
NowPlayinghandles AirPods, media keys and Control Center. Play and pause work as Space does, next and previous jump between markers, and seeking works from Now Playing. - Space anywhere:
SpaceBarCatcheris a window-wide key monitor, so Space plays even when nothing in the editor has focus. - Drag out: ⌥-drag a selection to Finder;
AudioFilePromisewrites the file after the drop. - On-device AI: speech-to-text uses
SpeechAnalyzer(macOS 26). Titles, summaries and sections use Foundation Models (Apple Intelligence).
Tooling
| Command | Does |
|---|---|
bin/build [--run] |
Debug build into build/Debug |
xcodegen generate |
Regenerate the project. Needed after adding files: project.yml globs folders. |
xcodebuild test -project MiniAudition.xcodeproj -scheme MiniAudition -destination 'platform=macOS' |
The test suite (it runs inside the app as its host) |
bin/screenshots |
DEBUG screenshot mode: stages scenes and captures the README images |
bin/profile <session> [--without part] [--analyzer] [--idle] |
Optimized build with PROFILING. Plays a session, attaches Time Profiler, prints a per-thread / per-area breakdown (bin/profile-report). --without stops one animated part to measure its cost by its absence. |
bin/notarize, bin/install |
Release DMG (signed, notarized, stapled) and installing it (see AGENTS.md) |
Conventions and pitfalls
- One feature per
EditorModel+*.swift. Keep files under about 300 lines and views small. Comments explain why, not what. - Layout:
- measured sizes never drive layout;
- no
ViewThatFits; - the inspector is a custom column, not
.inspector; - width-dependent layouts are picked from hidden copies at their ideal sizes.
- Layer-backed
NSViews setisFlipped. AppKit overrides a layer’s ownisGeometryFlipped. - SF Symbols: check a name with
NSImage(systemSymbolName:)before using it. On macOS 27, AppKit menu items needpreferredImageVisibility = .visibleto show their icons. - Undo in tests: set
groupsByEvent = falseand wrap registering calls inundoManager.step { }. Ending a group by hand hangs the test host.