PK/← All writing

ProgressCut · reliability note

I killed my Electron app twice to test recovery.

A local-first recorder should recover raw work after a process crash. I used SIGKILL during capture and real FFmpeg encoding to see what actually survived.

Source on GitHub ↗

I wanted a less flattering answer than “we have crash recovery.” What happens if the process disappears right now?

The question is not specific to Electron. It applies to a media exporter, downloader, migration runner or AI workflow that writes files locally. Restarting an app is easy. Recovering enough state to finish the user's work is the part worth testing.

I used an Electron recorder as the example below. The mechanics are simple: persist something, kill the process, restart, find the leftovers, rebuild, then inspect the result.

For ProgressCut, interruption can happen while screenshots are being captured or while FFmpeg is assembling the final MP4. In both cases, raw frames can exist while the output is partial or absent.

I started isolated Electron processes, created real session folders, sent SIGKILL, then started fresh Electron processes and asked them to find and rebuild what remained. SIGKILL ends a process immediately, so shutdown handlers cannot quietly clean up the state first.

The test found a rendering bug too: a story budgeted for 3.00 seconds encoded to 3.46 seconds.

Two interrupted ProgressCut sessions move through recovery manifests into rebuilt MP4 files
SIGKILL is a much more specific claim than “we have crash recovery.”

What had to survive

If there are captured frames, a later app process must be able to offer a rebuild without recording again.

The implementation is a per-session manifest stored under the user's local application-support directory. It contains the output folder, raw-frame folder, target duration, output format and timestamps. It does not contain screenshot data.

One manifest per session

A single active-session.json looks fine until two unfinished sessions exist and the second overwrite loses the first. Each session therefore gets a hash-keyed recovery record. Successful recovery removes only its own record.

Every manifest is written as a temporary file and then renamed into place.

Capture writes a per-session manifest, a crash preserves frames, and restart offers recover or keep frames
Temporary file → rename makes each manifest update atomic at the application level.
session A → recovery/<hash(A)>.json
session B → recovery/<hash(B)>.json

What I actually killed

The smoke test uses synthetic PNGs instead of recording a real screen. Everything after that is production code: Electron startup, session orchestration, manifest storage, story selection and FFmpeg.

ScenarioKill pointRestart must prove
Capture interruptionThird PNG persistedFrames are discovered and unchanged
Encoding interruptionReal FFmpeg has startedA fresh process finds and rebuilds the session

The parent launches each Electron app in its own process group and sends SIGKILL to that group. Killing Electron while leaving FFmpeg behind would not simulate an interruption cleanly.

After each kill, a fresh Electron process reads local recovery manifests. The harness rebuilds every discovered session, runs ffprobe — FFmpeg's inspection tool — on the output and verifies:

  • H.264 video;
  • yuv420p pixel format;
  • duration within 150 ms of the requested three seconds;
  • unchanged raw PNG bytes;
  • cleanup only for the session that completed.

The same test, in a different app

Replace “frames” with whatever your app leaves behind: uploaded chunks, edited documents, downloaded files, generated reports or queued jobs.

  1. Choose the smallest irreversible user value.
  2. Write a recoverable record before expensive work begins.
  3. Interrupt the real process at a deterministic checkpoint.
  4. Start a fresh process with the same storage.
  5. Verify the final artefact, then scope cleanup.
await runUntil("checkpoint:persisted");
killProcessGroup();

const restartedApp = await launchWithSameDataDirectory();
const pending = await restartedApp.discoverRecoverableWork();
const result = await restartedApp.resume(pending[0]);

expect(await inspect(result)).toMatchObject(expectedArtefact);
expect(await sourceBytes()).toEqual(bytesBeforeCrash);
expect(await pendingRecords()).toEqual([]);

I check cleanup separately because it is easy to get wrong: stale records cause repeated recovery prompts, while broad cleanup can erase another unfinished job.

The bug the crash test found

The first run failed on a stricter check than “the file exists.” A story budgeted for three seconds encoded to 3.458333 seconds. The concat demuxer repeated the final still frame to honour its duration, adding an unwanted tail.

A target of three seconds produced 3.46 seconds before a duration cap and 3.00 seconds after it
Playable output is not necessarily correct output.

I capped the encoded output at the story's calculated duration and kept a regression test that creates weighted moments and measures the resulting file with ffprobe.

"-movflags", "+faststart",
"-t", (story.totalDurationMs / 1000).toString(),
"-y", outputPath,

That is the whole reason to inspect the output rather than assert that a file exists. A report can exist with missing rows; a video can play and still have the wrong length.

What passed

On 2026-10-05, both SIGKILL scenarios recovered three synthetic frames on macOS 26.3 / arm64. The source PNG bytes were unchanged; rebuilt videos were H.264, yuv420p and 3.00 seconds long. An unsigned arm64 packaged app was also launched and exported through production IPC, with Sharp resolved from inside app.asar.

What this does not prove

This covers process crashes and nothing more.

  • Power loss or filesystem durability beyond atomic rename.
  • A multi-hour recording, sleep/wake, monitor changes or a full disk.
  • Screen Recording permission UX.
  • Developer ID signing, notarization, Gatekeeper or a clean-machine install.
  • Whether the story selector makes better summaries.

This is the boundary of the claim: process crashes are covered; the other cases are not.

Reproduce locallyProgressCut on GitHub ↗pnpm test:crash-recovery · pnpm test:packaged

Why I added this before release

For a local-first tool, useful data can already be on disk when the UI is gone. In this app, throwing away an unfinished session because FFmpeg died would be a bad default.

The test is a repeatable way to check the path:

Six steps: capture, kill, restart, discover, rebuild and probe
The test loop validates artefacts, not just whether the process starts again.

Next I need a two-to-five-hour real session with resource measurements, sleep/wake and an intentional interruption. This harness narrows the gap; it does not replace that run.