SCM Is a Free Mac App That Finds the Shot You Remember by Description, and Your Footage Never Leaves the Machine
Every editor knows the shot exists. Wide, rain on the window, someone at the desk, maybe the second shoot day, maybe the third. Knowing it exists has never been the problem. Finding it has, and that search still eats real hours on real jobs, usually at the worst possible moment in the schedule.
Most of the AI news in video is about making new shots. The more useful tool this week does the opposite. It finds the ones you already have.
SCM, short for Screen Memories, is a free, open-source Mac app that lets you search every photo and every frame of video in a folder by describing it. Type "two people at a kitchen table, morning light" and it returns the scenes that match, each tile showing a poster frame with a timecode badge. Open one, and the video jumps straight to that moment. Its own tagline sets the terms: local-first, with "no accounts, no cloud, no uploads."
The project appeared on GitHub on October 3, landed on Hacker News the next morning, and passed 160 points within two days. It carries the MIT license, so you can use it on paid work without asking anyone.
My position is simple. AI search belongs at the level of the drive, before you open an editing app, and SCM is the first free tool I have seen that puts it there.
Search that lives outside the edit
If you cut in Premiere, you may already know this idea. Adobe added Media Intelligence search in Premiere Pro 2025, which analyzes your footage "using on-device models" and lets you search clips by what is in them, as Larry Jordan explained in a July 2025 walkthrough. It works. It also lives inside one application, inside one subscription, and inside the media you have already imported into a project.
That last part matters more than it sounds. The shot you are hunting for is often in the footage you have not brought in yet: last year's shoot, the B-roll drive from the other camera operator, the stills a client sent over in a zip. An editor's real archive is a folder tree, not a project bin.
SCM works on folders. You import a folder (with ⌘I or by dragging it in), and the app starts watching it, picking up new files as they arrive and re-checking at every launch. The index lives under ~/Library/Application Support/scm by default, and the README says an environment variable called MEMORIES_DATA_DIR overrides that location, so in principle the index can sit wherever you choose.
So the search happens where the footage lives. That is the whole pitch, and it is a good one.
How it finds a shot from a sentence
Here is what SCM does with a video once you hand it over.
First, it looks for the cuts. The README says ffmpeg, the free video toolkit most editing software leans on somewhere, "scans each video for shot boundaries and builds a segment plan sampled to the density you pick in Settings." In plain terms: it breaks each clip into shots and decides how many frames from each shot to look at. More frames means slower indexing and finer search.
Then it looks at those frames with a vision model running on your Mac. The model does not write a caption for each shot. It turns each frame into a kind of numeric fingerprint of what it shows, and it does the same to your search phrase. Finding a shot means finding the fingerprints that sit closest to your words. That is why "moody hallway, flashlight" can find a shot that nobody ever tagged with any of those words.
You can choose among four vision models that ship with the app. The default is about 435 MB, the smallest about 214 MB, the largest about 850 MB. The one-time downloads (the vision model, Whisper's speech files, any extra text-reading languages, and the chat models if you switch that mode on) are the only network use; after that, in the README's words, "everything after that is offline."
Fingerprints are only one way to remember a shot, and SCM knows it. There are five search modes in all:
- Files finds whole photos and videos by what they show.
- Scenes finds moments inside videos and jumps to the timecode.
- OCR (optical character recognition) finds visible text: a street sign, a slate, a whiteboard, a product label. It uses Tesseract, a long-standing open-source text reader, in 36 languages.
- Dialogue finds what someone said, using Whisper, OpenAI's open speech-to-text model, running locally.
- LLMs, which you have to switch on, lets you ask questions of a local chat model that answers from the dialogue, text and filenames SCM has pulled out, with citations.
The dialogue mode deserves a closer look, because it shows someone thought about how editors actually remember lines. Results come in three tiers: "Exact line," for a phrase said in one go; "Exact words," for all your words inside one utterance or an eight-second window; and "Words spoken," for all the words anywhere in the same video. That third tier is the one you need when you remember the gist of a line but not the wording, which is most of the time.
One small feature saves a lot of grief: every file is identified by its contents rather than its name, so a clip someone renamed final_FINAL_v3.mov is recognized as the same clip you already indexed.
Why this matters more than another generator
A new video model gets you a shot you did not film. That is exciting and, for most paid work, still risky: licensing questions, consistency problems, a client who wants to know where the footage came from.
A search tool gets you back time on footage you own outright. Nobody argues about the rights to your own B-roll. Nobody asks you to disclose that you found a clip faster.
There is also the privacy point, which is not a small one for anyone working under an NDA. Cloud search tools ask you to upload footage, or at least frames from it, to someone else's server. SCM's README says "your media never leaves the machine," and the architecture backs that up: the models download once and then run on your Mac. For a documentary editor sitting on unreleased interviews, that is the difference between a tool you can use and one you cannot.
Put this into practice
Here is a first session that will tell you within an hour whether SCM earns a place in your kit.
Install it. On an Apple-silicon Mac running macOS 12 or later:
brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm
Start with one finished project, not your whole archive. Pick a job you know well, ideally one where you remember specific shots and lines. Import that project's media folder. You are testing the search against your own memory, so you need footage where you already know the right answers.
Pick the sampling density deliberately. In Settings, Video Search, choose how densely SCM samples each shot. Start in the middle. If it misses short inserts later, raise it and re-index that folder.
Run three kinds of search. Describe a visual moment you remember in Scenes mode. Search for words you remember from a line in Dialogue mode and see which tier it lands in. Search for a piece of visible text, a sign or a slate, in OCR mode. Note which mode finds things fastest for the way you remember footage.
Write searches like a shot description, not a keyword. "Close-up, hands, coffee cup, warm light" beats "coffee." The vision model compares your whole phrase to what it sees, so composition and light words help.
Then widen it. If the first project went well, add your stock and B-roll folders. Those are the drives where search saves the most time, because nobody ever logged them properly.
Where it falls short
It is Mac only. The app is built for macOS. Windows and Linux editors are out for now.
Indexing a big library will take time, and nobody says how long. The README gives no speed figures. A few terabytes of footage at a fine sampling density could tie up your machine for a long while. Index overnight and start small.
It finds frames, not edits. Results open the video at the timecode. There is no panel inside Premiere, Resolve or Final Cut, and no way described in the README to send a found clip straight to a timeline. You still note the timecode or drag the file in yourself.
Description search has a ceiling. The vision model knows general things well: rooms, weather, people, objects, light. It knows nothing about your project's private vocabulary. It cannot tell your lead actor from your second lead by name, and "the take where she nails it" means nothing to it. That is where the dialogue and text modes earn their keep.
It is brand new. The repository is days old. Expect rough edges, and do not let it become the only record of where your footage lives.
The format list is not documented. The README covers videos, photos and GIFs, but it does not list which camera formats it reads. Test your actual camera originals, especially anything raw or log-encoded, before you trust it with a whole job.
What to do with it
The interesting thing about SCM is how unglamorous it is. It does not make anything. It does the job assistant editors used to do with a notepad and a long afternoon, and it does it without sending a single frame anywhere.
Generators will keep getting the headlines. Tools like this decide whether you go home on time.
So give it one real project this week. Load footage you know by heart, search for five shots you remember, and count how many it finds and how fast. If it finds four, point it at the B-roll drive you have been avoiding. If it finds one, you have lost an hour and learned exactly where description search breaks for your kind of footage, which is worth knowing before a client deadline makes you find out.
Medium metadata
Title: SCM Is a Free Mac App That Finds the Shot You Remember by Description, and Your Footage Never Leaves the Machine
Subtitle: An open-source, offline search for every frame, line of dialogue and on-screen word in your video folders. How it works, how to test it on a real project, and where it falls short.
Tags: Video Editing, Filmmaking, AI Tools, Open Source, Mac
Estimated read time: 8 minutes