Skip to content
All posts

Building AtomCut

4,555 free sound effects, none loud enough to hurt you

A CC0 sound-effects archive inside AtomCut's browser video editor: 4,555 sounds, waveforms before you download, and an ear guard that measures loudness.

SShayanBuilding AtomCut

The Sound Effects panel in AtomCut: a dense list of interface sounds with real waveforms and durations, one row playing mid-sweep, and a rail of packs with their counts

AtomCut ships a sound-effects archive: 4,555 public-domain sounds from 37 publishers, searchable inside the editor, auditioned one row at a time, dropped onto a lane as ordinary audio clips. Clicks, impacts, whooshes, glitches, foley, jingles, voice. Every one is CC0, so whatever you make with them carries no conditions. AtomCut is a free motion design and video editor that runs in the browser, with no account and no install.

The interesting constraint was not finding the sounds. It was that browsing them is physically unpleasant, and fixing that properly needs a number nobody publishes.

Where the 4,555 sounds come from

It started at seven hundred, almost all from Kenney, who has put game assets into the public domain for over a decade. Nine audio packs, CC0, no attribution required.

Those packs are game audio, though, and a motion piece reaches for things a game rarely needs: a pass-by whoosh leading a cut, a glitch to cut through, a riser that ends on the beat, a sub hit under a title. So the archive carries a second pack synthesized from DSP written for AtomCut. Filtered noise bursts, gated bit-crushed stutters, two-note stingers. That keeps it reproducible from a script and unambiguously ours to give away.

Seven hundred is a demo, not a library. The rest came from OpenGameArt, which has 831 sound-effect submissions filed under CC0 — seventy-nine of them are here, taken most-downloaded-first, because a community's own download count is a judgement about quality that no heuristic in a script could make.

Two things about that were not optional. The licence is read off each submission's own page, never inferred from the search filter that found it — same claim, same database, but it is the one fact in the whole pipeline that cannot be recomputed later if it is wrong, and a mislabelled file in a public archive is not something you fix with a patch. And the same sound is genuinely in a dozen compilations, so source bytes are hashed before transcoding and the first pack to ship a sound keeps it. One 51-sound pack turned out to be 51 duplicates of another and was dropped whole. Content-addressed storage would have deduplicated the bytes and left all eleven identical rows on screen: that fixes the bill, not the library.

What is not optional and was a surprise: hundreds of page fetches from one community site looks exactly like abuse from the far end. Forty packs in, every request started failing at the socket — not a 429 with a Retry-After, just a refused connection, which no status check would have caught. A crawl needs both halves: a pause between requests to stay inside what the site would have allowed anyway, and a backoff to get home when the pause was not enough.

The Sound Effects panel: packs as cards, each with its sound count, size and publisher, and a rail of categories on the left
Ninety-one packs, deepest first. Every count, size, blurb and tag on this screen was derived by the pipeline. None of it was typed.

Nothing in that screenshot was written by hand. A pack arrives as a zip of files called forceField_000.ogg and back_001.ogg, and hand-tagging a few thousand of those is a week of work that goes stale the next time a pack updates. Names, categories and tags are derived from the path by a taxonomy that lives in the app's core, which is the same code that runs if you publish your own pack.

One detail there was sharper than expected. Take numbers cannot be read off a single filename. Kenney ships back_001, back_002, back_003 in one pack and drop_000, drop_001 in another: one-based in one place, zero-based in the other, sometimes in the same download. Guessing per file produced a shelf that began at "Back 2". Whether upstream counted from zero is a property of the series, not of any file in it, so naming reads the whole folder at once. It also drops the number when it distinguishes nothing. There is one bong in the pack, so it is called "Bong".

Why peak is the wrong thing to measure

Browsing a sound library on headphones is mildly hostile. The files were mastered by hundreds of different people, the levels are all over the place, and you set your volume for whatever quiet thing you last clicked. Then you click a brickwalled impact. "Turn it down first" is useless advice, because you cannot know which file it was until you have already heard it.

So sounds play through a ceiling. Anything above it comes down to it, anything below is left alone.

The obvious way to build that is to measure each sound's peak and scale it. It barely helps. A heavy sustained impact and a 3 ms interface tick can both peak at exactly 1.0 and sit fifteen decibels apart to a human ear. Peak is not what hurts; sustained energy is.

So the pipeline measures perceived loudness instead and publishes one number per sound, stored beside the waveform in the manifest. It has to be known before the audio downloads, because the guard decides how far to duck a sound the moment you first click it.

The first version measured mean energy over the audible part of the file. A test written to prove the point disagreed:

expect(loudness(oneSecondTone)).toBeGreaterThan(loudness(threeMsClick) + 6)
  → -3.01  is not greater than  2.99

They measured identically. Both were full-scale sine, so both had the same mean square over their own length, and dividing by "their own length" is exactly the mistake. The ear does not judge a 3 ms transient and a one-second tone on equal footing. It sums energy over roughly 200 ms.

The fix deleted more code than it added. Integrate over the audible span, then divide by at least the ear's integration window:

js
const window = Math.max(audibleSamples, 0.2 * sampleRate);
return 10 * Math.log10(sum / window);

A 3 ms click now measures about 18 dB under a sustained tone of the same peak, which is roughly what it sounds like. A block-gating pass written earlier to handle long silent tails turned out to be unnecessary, because trimming to the audible span already did that job.

The second decision matters as much: this is a ceiling, not a normaliser. Normalising to a target would make every sound equally loud and destroy the one thing you are auditioning for, which is whether this impact hits harder than that one. Only what is over the line moves, and only down to the line. When the guard is holding something down, the player says so, because a control that quietly changes what you hear while you judge a sound has to admit it.

A whoosh playing: the row is highlighted with its waveform swept, and the player below shows the name, 279ms, hit 22ms, -12.2 dB and a shield chip reading minus 6
The chip beside the name says the guard is holding this one 6 dB down. It measured −12.2 dB against an −18 dB ceiling.

A list, not a grid

Artwork is scanned in parallel: a grid of thumbnails, eight at a glance. Sound is read in series. You play one, then the next, then back to the one before, holding a queue in your head the whole time. That is the one thing that justifies a different surface from the Library's grid, and it is why the keyboard does the work here.

Space plays. ↑↓ walk the shelf and play as they go. ←→ scrub. ↵ drops the sound on the playhead. Space plays a sound here and never your composition, which needed no new machinery: the app already had one opt-in attribute meaning "this surface owns raw keystrokes", and the window wears it.

Underneath, none of this is a second catalogue. It reuses the Library's client, manifests, matcher and stars, because that file already opens by explaining that "library" once meant five different things in this app at once, with five card designs and five ideas of what a preview is. A sixth would have been the mistake it exists to have ended.

The Whooshes category showing six whoosh sounds drawn from two different packs, each with a waveform and duration
A category gathers one kind of sound from across the archive. Opening this shelf downloaded only the packs that are mostly whooshes — not every pack with a whoosh in it.

What loads when you open the panel

67 KB. One index, a row per pack. That is all — for 4,555 sounds across 91 packs.

The landing lists packs, not sounds. Opening a pack fetches that pack. A category reads the index to work out which packs it should fetch and fetches only those. A search does the same against a list of every searchable word per pack, so whoosh reaches a handful and ignores the rest. Audio arrives only when you press play.

The design was written on the claim that the archive could grow sixfold without the panel getting slower. Then it did, which is a rarer thing than it sounds: most scalability arguments are never tested, they are just believed until something else changes. This one was mostly right and wrong in one specific place.

A pack's row in the index listed which tags it contained. At twelve curated packs, a tag named what a pack was about. At a couple of hundred general ones, almost every pack contains at least one of everything — one stray impact in somebody's zombie pack — so clicking Impacts asked for almost every manifest in the catalogue. The fix is not a cap or a cache: the row carries how many of each tag now, not which, so a shelf fetches the packs the subject actually belongs to and the rail can count sounds instead of packs. One stored fact, and the wrong one had been sitting there since the day the catalogue had twelve packs in it.

A preloader sits on top, deliberately timid. It warms the next sound after you play one, leans a little further ahead the longer you keep going, never more than four. It skips anything long, holds the warmed set inside a memory budget, evicts the oldest, and fetches strictly one at a time. Preloading is a convenience that must never become the reason a tab is using 800 MB.

Writing that down exposed a bug shipped months ago. The Library browser called the catalogue's search with an empty query whenever you opened it, and an empty query matches every pack, so it downloaded every manifest there is. At sixteen sounds nobody noticed. At seven hundred it was most of a megabyte, fetched to draw a number next to the word "Sounds". The rule it broke was already written down in a comment in the file next door: an online sound is not a kind of catalogue entry, it is a search result.

Putting one on the timeline

Drag a row onto the stage or a lane, or press ↵ to drop it at the playhead. The bytes stream in behind the gesture, so nothing waits on a download.

By default the clip starts slightly early so the sound strikes on the playhead instead of beginning there. Every sound carries a detected transient, measured at publish time, and that offset is the difference between a whoosh that lands on the cut and one you spend thirty seconds nudging.

Three audio clips on the AtomCut timeline: Riser Tension across three seconds on track one, a whoosh on track two, and Transition Glitch selected at the playhead, with the inspector showing volume, fades, the sound effect stack and role SFX
Three sounds placed with Enter. They arrive as ordinary audio clips, tagged SFX, with fades, an effect stack and keyframable gain.

What it does not do

The landing settings (hit-on-playhead, fades, gain, pitch) apply when you click, press ↵ or use Add. Dragging a sound onto a lane places the file as it is. Making the two agree means teaching the generic media drop handler about a panel preference, which would also apply it to a music track dragged out of the media library, and nobody asked for that.

Freesound is written but not switched on. Its CC0 subset is about 380,000 sounds — the only source at that scale that is genuinely ours to redistribute — and the adapter is in the registry waiting on a free API key. Past roughly five thousand sounds the audio should also stop living in a git repository and go straight to storage with only the manifests committed, which is a change to where the bytes sit and not to the pipeline, because its stages already meet at a directory of files.

Open the Sound Effects panel and hold ↓.