PK/← All writing

Open source · algorithm deep dive

From 10,000 Frames to 30 Moments

Two hours of coding generates ~14,000 screenshots, most of which are the cursor blinking. Cutting that to 30 moments — no ML, no cloud, a few seconds on your laptop — takes five passes over the same data, each killing a different category of noise.

Read on DEV ↗Source on GitHub ↗

Two hours at 2fps is around 14,000 frames. Most of them are noise: the cursor blinks, the system clock ticks, nothing in the code moved. The first step isn't clever. It's just getting rid of those.

Five passes, five different cuts

Each stage kills a different category of useless frame. Real numbers from a 2-hour session:

Five-stage pipeline: 10,847 raw frames → dHash 1,203 → SSIM 284 → segments 48 → story 30 moments
Frame counts from a real 2-hour session

Stage 1 — dHash: kill the identical frames

89% gone in one pass. If you record at 2fps and type at a normal pace, maybe one keystroke lands per 3–5 frames. Everything in between is the same screenshot with the cursor in a different blink state.

dHash resizes each frame to 9×8 pixels, converts to greyscale, then compares each pixel to the one on its right (1 if brighter, 0 otherwise). 8 rows × 8 comparisons = 64 bits per frame, stored as a BigInt. XOR two hashes and count the differing bits. Distance ≤ 8: duplicate.

dHash pipeline: source frame to 9x8 greyscale grid to 64-bit comparison bits, Hamming distance decision
1920×1080 → 9×8 greyscale → 64 gradient bits → Hamming distance

Counting those differing bits naively loops 64 times. Kernighan's trick loops once per set bit, so for near-duplicates that's typically 2–6 iterations, not 64:

export function hammingDistance(a: DHash, b: DHash): number {
  let diff  = a ^ b;      // XOR: 1 where bits differ
  let count = 0;
  while (diff > 0n) {
    diff &= diff - 1n;    // clears the lowest set bit
    count++;
  }
  return count;
}

Stage 2 — SSIM: kill the boring frames

Unique isn't the same as interesting. After dHash you still have ~1,200 frames that each look slightly different, but most of that difference is one more character on a line. You need a score that reflects how much the visual state actually changed, not just whether the pixels moved.

SSIM (Structural Similarity Index) compares luminance, contrast, and spatial structure between consecutive frames. Score close to 1: nothing happened. Score close to 0: something worth keeping.

Novelty score timeline showing spikes when significant visual changes occur, with selected moments marked as purple dots
SSIM novelty over time — purple dots are the frames that made the cut

Full-frame SSIM would average out a big change in one corner against stagnation everywhere else. Instead the frame is tiled into 16 regions, SSIM computed per region, then averaged. A file-save that changes the status bar doesn't steal credit from a tab switch that changed everything.

The formula needs mean, variance, and covariance for each region. The naive path does two loops: one for the mean, one for the variance. Both collapse into one using Var(X) = E[X²] − E[X]²: accumulate Σx, Σx², Σy, Σy², Σxy in a single pass and derive everything after:

let sumA=0, sumB=0, sumAA=0, sumBB=0, sumAB=0;
const n = (x1-x0) * (y1-y0);

for (let y=y0; y<y1; y++) {
  for (let x=x0; x<x1; x++) {
    const pa = a.pixels[y*W+x] ?? 0;
    const pb = b.pixels[y*W+x] ?? 0;
    sumA  += pa;    sumB  += pb;
    sumAA += pa*pa; sumBB += pb*pb; sumAB += pa*pb;
  }
}

const muA = sumA/n, muB = sumB/n;
const varA  = sumAA/n - muA*muA;   // Var(X) = E[X²] − E[X]²
const varB  = sumBB/n - muB*muB;
const covAB = sumAB/n - muA*muB;

Stage 3 — Segmentation: budget the moments fairly

Sessions aren't uniformly active. There's 20 minutes of flow where you're actually building something, then 10 minutes reading docs, then another sprint. If you spread the 30-moment budget evenly across the timeline, the reading gaps eat slots that should go to the interesting parts.

Temporal segmentation showing timeline divided into active and idle windows with budget bars below
Dense segments get proportionally more of the 30-moment budget

The pipeline splits the frame sequence into activity windows by novelty density. Each window's share of the budget scales with how much happened inside it. A 20-minute sprint might take 8 moments; a 10-minute idle stretch gets 1.

What comes out

30 moments → a GIF or MP4, under 30 seconds, no input from you. Each moment holds for a duration proportional to the activity level of its segment: busier parts play faster, idle parts barely appear.

Compression ratio: 10,847 grey frames in the top bar, 30 purple moments in the bottom bar
10,847 raw frames → 30 moments; bar widths are per-moment hold durations

No frames leave the machine. The app calls no model. Run it twice on the same input and you get the same output. That last part matters more than it sounds: a non-deterministic story compressor is harder to trust and impossible to debug.

What's next

Pixel distance is one way to measure novelty, but SSIM can't tell “opened a new file” from “typed one character” because both move roughly the same number of pixels. A CLIP embedding delta would catch that difference. That's the next scorer to try, once there's a feedback loop to check whether it actually picks better moments or just different ones.

The feedback loop needs data first: a drag-to-remove UI so users can mark which moments they'd cut. A few hundred sessions of that and there's something real to validate against.

The rest of the pipeline doesn't change. The novelty scorer sits behind a port in packages/engine: swap it out, everything else stays.

Code

MIT, pnpm monorepo, 187 tests, CI on macos-latest. The algorithm lives in packages/engine/src/ with no Electron dependency, runs in plain Node. If you want to poke at the frame selection logic without building the whole app, that's the entry point.

PK
Pavel Kazantsev

AI/ML engineer building evaluation systems, research pipelines and evidence-backed AI products.