7 ms·
Indexing a year of video locally on a 2021 MacBook with Gemma4-31B (50GB swap)
- andai 5mo agoAwesome. Say, this is very comprehensive. I was vaguely aware of all these pieces existing (except for running a facial recognition database at home o_o), but it's really neat to put them all together like that.
- asenna 5mo agoThanks! I was honestly casually trying it out on the side with Claude's help. And I was actually pleasantly surprised to see how good the result was. Still blows my mind I can do all this from my 2021 MBP. I'll try to do a post once I have the next steps working (helping with planning and editing videos with Davinci Resolve).
- ahknight 5mo agoI also have a 64GB M1 Max and am similarly impressed with what that workhorse can do. The M5 tempted me -- a lot -- but then I looked at what I was already getting done on that machine and just couldn't justify it ... yet. Someday, surely, but not yet. Gemma4 gave all my local projects new life, just like what you did here. Great job. Long live the M1 Max!
- asenna 5mo ago100% Although knowing how good these local models are getting, I am now eyeing the upcoming M5 Ultra Mac Studio (256gigs perhaps). But knowing how crazy the market is, it might be a year before I get the chance to get my hands on it. If it even launches by WWDC.
- throwa356262 5mo agoI ran Gemma on a 2015 thinkpad to do something similar. Fortunately, I could upgrade the memory otherwise it would have been a painful exercise. Not gonna lie, llama.cpp had the fans spinning at max speed. But it worked and I got the job done.
- iMerNibor 5mo ago> the fans spinning at max speed This always confuses me - don't people want their computations to run as fast as possible and thus inevitably produce more heat that needs to be vented? I suppose sometimes it is just an analogy for "its utilizing 100% of my resources" (which I'm guessing it is here), but I've definitely had people say it as an actual complaint in different contexts
- 0xbadcafebee 5mo agoFans shouldn't be running at max speed if the model fits in RAM with room to spare for context. Usually fans max out when the model doesn't fit and the CPU is chugging to make up the difference (or the user didn't tune LLM settings)
- mmmlinux 5mo agoI didn't realize GPUs operate with zero heat output.
- dist-epoch 5mo agoWhat people complain is when they visit a blog with two images and the fans are spinning at max speed because the blog has 100 trackers.
- overfeed 5mo ago> I've definitely had people say it as an actual complaint in different contexts I think fan loudness is an outgrowth of conspicuous consumption because a certain OEM decided to make it a marketing bullet-point. I was equally disappointed by by people - especially device reviewers - banging on the drum that phones made of plastic "didn't feel premium", and we got phones with glass backs that have to be shoved into plastic cases (because plastic is the near-perfect material to protect fragile phones screens and innards)
- evilduck 5mo agoWow, someone mentions fan speeds on their Thinkpad and you go into an unprompted dog whistle rant mode. A certain OEM lives rent free in a blighted mind.
- egorfine 5mo ago> generative AI video has no place on a real travel brand I am pretty sure that the vast majority of Airbnb hosts would not agree with you. > equals TripAdvisor crucifixion I have no idea how the Airbnb hosts with fake listings survive, really.
- asenna 5mo agoHaha. It's honestly something that I've been struggling with myself. I'm running this safari lodge but I don't want to go down that route of slop videos! But on the other hand, genuine videos do take time and slows down the process.
- desro 5mo ago> The skill is open at ~/.claude/skills/video-index/. If you're working on something similar (indexing personal archives, getting a local model to do real archival work, building agents that drive editing tools), I'd be glad to compare notes. When your Claude wrote this post they might not have selected the right URL to share, unless your home folder is exposed. Care to share the skill files?
- asenna 5mo agoOops! My bad. Fixing it now. And yeah, I can share the Skill file. Give me 5 mins.
- asenna 5mo agoOk I scrambled to finalize a name for it and create a new repo for it - https://github.com/Simbastack-hq/framedex https://github.com/Simbastack-hq/framedex PS - I just put this together in the last few mins, removed my personal files and references. So it's not tested properly, please let me know if any issues. It's still an early hack, but I have thousands of still images as well from my camera which I've not processed and I need to do the same analysis for those. So I'll continue working on it, but happy to receive any PRs if anyone finds any use for it. I'm tired of having a backlog of thousands of images and videos, leaving it for later.
- jaggederest 5mo agoHey friend, try something in this ballpark, your post has a bunch of painful AI tropes: https://github.com/blader/humanizer https://github.com/blader/humanizer You get a pass here because you're doing really cool stuff but it's kinda tough to read past the AI nonsense, and it's relatively easy to screen out "it's not x it's y" kind of things and the bolded bullet points.
- asenna 5mo agoThanks for this! This is exactly what I was looking for. Tbh, I have a lot of thoughts and ideas and things to share and I do spend time and effort trying to de-AI-ing it but this should help a lot. I'll try it out. In fact, I was expecting getting shit on by HN readers for this but was pleasantly surprised that readers moved past it.
- egorfine 5mo agoThanks for the article! I have a beefy M5 Pro and I'm eagerly looking around for ways to use local models (specifically Gemma4 & Qwen3.6). This is an excellent thing to do. Especially that LLMs excel at batching thus you can index multiple photos and videos in parallel for no performance penalty.
- busfahrer 5mo agoI have been contemplating a M5 Pro MBP, but for the life for me I wasn't able to find benchmarks for real-world models, do you happen to know how many tokens per second roughly you get with MoE models like Qwen 3.6 35B/A3B or Gemma 4 26B?
- ahknight 5mo agoI'm not normally one to share videos as answers, but this particular fellow does a LOT of work with local AIs and Macs and happens to have a nuanced answer. https://youtu.be/XGe7ldwFLSE https://youtu.be/XGe7ldwFLSE
- egorfine 5mo agoQwen 3.6 35B running on oMLX 0.3.9rc1: on oMLX I get 86 t/s on Q4 and 74 t/s on Q6. Bear in mind that ttft on MLX is much much faster on M5 Pro as compared to M4 Pro. Also bear in mind that those figures are with NO optimizations whatsoever: no MCP, no DFlash. I am waiting for both to be released for the Qwen models.
- herf 5mo agoTwo questions: 1. What is the search index? 2. The "description.md" example has things like "faces -> cluster_id". Is this from Davinci Resolve's face index? Things like faces+names and locations are really important with photo collections, but general LLMs don't handle them so well.
- asenna 5mo ago1) It's just simple plain-text `.description.md` sidecar files, one per clip, sitting next to each video. Something which I can query later - Like when brainstorming with Claude "I wanna make some videos of the Luxury rooms in the lodge" and it knows what all videos could help here (going through the files). There's also a folder root level files that aggregates the text descriptions to make it easier to find. I've just attached an image in the blog showing an example - https://blog.simbastack.com/_media/gvcycx2n.png https://blog.simbastack.com/_media/gvcycx2n.png 2) No - nothing from DaVinci Resolve. Framedex is a standalone pipeline. Resolve isn't involved. Faces come from insightface (the open-source buffalo_l pack - RetinaFace for detection), running locally on CPU. For each clip it detects faces in the sampled frames, embeds them, and writes rows to ~/.framedex/faces.db. Tbh, this part I know it's building up in my local DB but I haven't tested how good is it. Will check them out properly soon. But yeah, on your broader point that's why framedex deliberately does not ask the LLM to handle faces or locations. ---- Faces → insightface / ArcFace embeddings. Deterministic, comparable across clips. The vision model only contributes a rough people_count; it never tries to identify anyone. Locations → EXIF GPS via exiftool, reverse-geocoded through Nominatim/OpenStreetMap. Hard metadata, not a guess. The LLM only does what it's good at: scene description, mood, shot type, keywords, keep/review/cull rating (this last part is also debatable though).
- asenna 5mo agoUPDATE: Quickly created a repo for this - https://github.com/Simbastack-hq/framedex https://github.com/Simbastack-hq/framedex (MIT License) It's not tested properly after I genericized it. Will try to go through it properly and add more updates. Two big things on my TODO: 1) Make use of this indexing and using Claude's help, make video editing faster with Davinci Resolve (now that I have a good index of all the content) 2) I currently did this for videos, but I want to add more things to this for my thousands of still images of my camera - need to make sense of them. So I'll be working on this as well.
- maxothex 5mo ago[flagged]
- gitowiec 5mo agoReading this text feels strange, sentences seems to be detached
- cataphract 5mo agoI had exactly the same impression, and I recall seeing this style other times recently. First time I thought it was just bad writing skills, now I'm thinking it's AI generated.
- asenna 5mo agoI'm the author, yes it is AI-assisted. You can make AI-generated content without it being slop. Slop, to me at least, is content that's wrong, padded, or generic. I see the cadence / short-sentence issues but if there's something else beyond those, I'd actually want to know what made it feel bad. I would've put off documenting what I did over the weekend but instead, I did document everything, spent quite some time (several iterations) and effort to make sure it does not hallucinate and writes in my own tone and voice. I'm sure it could be better but the content is not made-up. At a time where most of us software engineers have changed our workflows to let AI write 80+% of our code using agents, I feel writing is heading the same way. It then becomes a matter of taste, whether it's done well or not. If you're looking clues and signs for whether a content has used AI, you're going to be disappointed over the next 12 months. If it feels jarring right now, I'll work harder on the workflow so it feels more natural next time (someone shared this project with me - https://github.com/blader/humanizer https://github.com/blader/humanizer). But this clearly allows me to make content which I wouldn't have done earlier.
- cataphract 5mo agoI'm not philosophically against AI or anything, but I think this needed some heavy editing. I did not even initially think upon seeing this style for the first time that it was AI-written, because I would associate AI-written text as fluffy. This staccato instead looks like the model was told to be terse and informal. I think the informality doesn't help either -- it's not that you can't have a well-written colloquial text, but I think it's harder to pull off. Here is an example: > Gemma returned people_count: "many" instead of an integer. My vision prompt literally said integer or the string "many" if >10. Gemma followed instructions correctly; the bug was schema design. The fix was a stricter prompt (integer 0-99 with explicit guidance to estimate) plus a coercion in the parser for the legacy "many" responses. Don't union-type schema fields. Pick always-int or always-string, never "int or this one specific string," because every downstream consumer pays for the choice.
- brcmthrowaway 5mo agoSo do they run the lodge or what?
- theodorewiles 5mo agoMy take is that B2C AI applications are kind of structurally limited by how hard it is to build personalized context. The idea of capable local models could be a huge unlock here if they are able to do the bottom-up context collection research / tagging / etc. at scale.
- enos_feedler 5mo agoIs it really local models that unlock this? Surely stateless model APIs would yield the same benefits? I get that local can be “cheaper” depending on usage, but we’ve been renting storage and compute from clouds at a premium for ages..
- asenna 5mo agoA huge thing here was the massive amount of data that was just processed - I went through about 1TB of files over 24 hours. Using API to analyze even a subset of this would've been painful imo.
- enos_feedler 5mo agoI thought about that in this video case and it's true. I thought the parent comment was making a broader statement about local models in general. But even with video, if it was stored in private cloud storage near the LLM could this still have worked efficiently? What are the most painful elements of this whole setup / work environment if everything was cloud?
- asenna 5mo agoOh yes, if everything is cloud, then this is a non-issue. The few other points of consideration would be: 1) Cost - I was considering using Sonnet for this but there's always the concern of reaching limits OR the API cost if you're using the API. The feeling of knowing you have a capable model in your hands without any limits is actually pretty awesome. Your mind starts running at what else can I throw at it to do grunt work. 2) Privacy issues - same as with moving to cloud. 3) Reliability issues - I know from experience Claude uptime has been pretty bad the past few months 4) Restrictions - Claude has been pretty heavy handed with their restrictions lately, anything which remotely triggers there flags gets an instant denial (or worse, an account ban). Often these are false-positives. I love the value I get from Claude but there's a different kind of freedom you get with local, capable models.
- zazibar 5mo agoThe subject matter is interesting but the amount of slop makes it difficult to read through. Yeah, it's great that you can throw your technical problems at Claude without caring much about the generated output but treating your own writing that you actually want to share with the world the same way is a terrible idea.
- asenna 5mo agoTbh, I did spend a lot of time trying to ground it and de-slopify it - verified nothing was halucinated and went through 10 iterations to get to this. It's almost like wrestling with Claude and I knew it would be tough on HN. But because of the fear of non-perfection, I used to put away things like creating this article or even posting it anywhere. And I do think the article has real value that HN would appreciate (I am myself an HN-enthusiast). I'll try more. Someone else shared this project which would be really helpful - https://github.com/blader/humanizer https://github.com/blader/humanizer Also a side note, the blog is posted on my self-created Slopit.io platform which is purely meant for your personal agents (working along with you) to post content - I recommend trying it out. https://blog.slopit.io/this-blog-post-is-slop/ https://blog.slopit.io/this-blog-post-is-slop/ I know, things are getting difficult with all the slop around, but my personal opinion is, as the agents get better at writing, the "annoying-ness" factor reduces and pieces of substance will still be appreciated, even if it was written by agents. This and the fact that agents aren't going away. If I've automated a lot of my coding, I feel like engineers like me would naturally progress to also taking agents' help to write useful content. PS - this comment was 100% hand-typed.
- teach 5mo agoFor what it's worth, I really enjoyed this read and almost came here to comment "this is the most enjoyable llm-assisted article I've read in a while" The tells were unmistakable but it still had a human touch, so I for one am glad you published anyway.
- asenna 5mo agoI'm definitely learning and hope to do better next time but your comment truly means a lot. I kid you not, I've taken a screenshot of this to motivate me next time I'm doubting publishing :)
- cold_harbor 5mo agothe reason 50GB swap is even viable here is Apple Silicon's memory bandwidth. on x86 that much swap would make inference unusably slow
- throwawaytea 5mo agoMemory bandwidth or storage bandwidth?
- bahmboo 5mo agopotato potatoh
- yardie 5mo agoNow I have another project for this weekend! I also have tons of video and not a lot of time to index them.
- ngai_aku 5mo agoI’d like to do something like this for the collection of home videos I have piling up, but I’m still on 16GB M1. Any hope of getting decent results with smaller models? If not, does anyone have tips on GPU rental? I have a Claude max sub and plenty of OpenRouter credit, but I don’t feel good about uploading my family’s private videos
- oceanus 5mo ago[flagged]
- dang 5mo agoCould you please not post generated comments to HN? It's not allowed here. See https://news.ycombinator.com/newsguidelines.html#generated https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 https://news.ycombinator.com/item?id=47340079. We ban accounts that do this and I don't want to ban you, so please write everything that you post to HN by hand. Of course, it's impossible to know for sure what was LLM processed or not, but we're getting complaints about some of your posts and, upon inspection, the complaints seem justified.
- clueless 5mo agoThis sounds like a great capability to be added to immich
- asixicle 5mo agoOr Stash lol
- Confiks 5mo agoI'm not quite sure why all that swapping is necessary. I really does age your SSD quite fast considering the enormous memory bandwidth required. Gemma 4 31B at 4-bit quantization should only be around 19 GiB [1], not 28.4 GiB. I'm not feeding it images regularly, so I'm not sure how much memory it needs to get those into context, but I can't imagine it is more than 10 GiB. The activity monitor does show all kinds of Electron apps active, on top of a presumably model-loaded Handy and a virtual machine for Claude Code, so I guess that's the real root cause for all the swapping. If your laptop starts trashing I can't imagine you have any use for those apps, which will grind to a halt. [1] https://huggingface.co/mlx-community/gemma-4-31b-it-4bit https://huggingface.co/mlx-community/gemma-4-31b-it-4bit
- asenna 5mo agoYeah to be fair, I could've cleaned everything up but this was taken when I was doing other work on my laptop while the screenshot was taken. Although slightly laggy, I was impressed by the fact that I was still able to work on other things and have a bunch of tabs open on my Brave browser.
- genxy 5mo agoWhy did you destroy your own voice to have it replaced by AI ?
- mainaisakyuhoon 5mo agoI really struggled to read the AI slop in this.
- carpo 5mo agoThis is great. I wish I had enough ram for a local model. I just spent the last few weeks writing something very similar, but I made it a local Electron app with Whisper, ffmpeg and I added semantic search and embeddings for chatting with the videos. It talks to Claude for the vision analysis, tagging and video chat. Do you only send one image for yours? I used a customised scene detection algorithm to find multiple different images per video and then send them all in one request to Claude (along with the subtitles). It's definitely the most expensive part. Using Sonnet 4.6 for the analysis and Haiku for the tagging costs about $1 for an hour of footage, I can imagine it would be slow locally.
- nl 5mo agoTry some of the models on OpenRouter if you are looking to save money. Gemma 4 31B is $0.12/M input, $0.37/M output vs $1/M input, $5/M output for Haiku. There are other options that are good too. Gemini 3.1 Flash Lite is great for this kind of thing (NOT Gemini 3.5 Flash though - the pricing for that is bad). https://openrouter.ai/google/gemma-4-31b-it https://openrouter.ai/google/gemma-4-31b-it
- carpo 5mo agoCheers, I'll give it a try. How are those models at returning structured results? When I was writing the prompts for the analysis step and testing with older Claude models, it would have trouble structuring the XML consistently. Sonnet 4.6 handles it really well.
- nl 5mo agoUse function calling/tool use, not XML output. The models are all trained for that now. Ie, instead of telling it to generate <name>Name</name> <age>19</name> <address>whatever</name> give it a function details(name: string, age: int, address: string) That is actually a JSON schema, and the models do great at it. Here's the claude docs, but they are all similar: https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools https://platform.claude.com/docs/en/agents-and-tools/tool-us...
- pavlov 5mo agoThe content is good, but this LLM writing style gets tiresome. Everything is a revelation: >“I bought it for Chrome. It's running a model that didn't exist when I bought it.” Well duh, personal computers run new software. That’s literally the whole point. The Apple II didn’t sell on the strength of the preinstalled apps.
- asenna 5mo agoAuthor here. I totally hear you. I wasn't expecting this to do well on HN for exactly this reason. But I've mentioned elsewhere - if it wasn't for all the AI-assistance, I would've put-off documenting everything that I did and not even get to the writing part. But yeah, I'll be working on the workflow to make the next write-up better, more humanized.
- moinism 5mo ago> Every AI video editor on the market assumes your footage is already labeled Shameless plug: I'm the founder of Chat Octopus, an AI media assistant, and it actually 'looks' at the videos to understand them before creating a cut.
- coldtea 5mo agoThe post is a mix of human and AI writing and the AI-mannerisms get on the nerves. At least it has a clear topic and some actionable insights and code examples.
- danborn26 5mo ago[dead]
- RealMarcus_AI 5mo ago[flagged]
- dwa3592 5mo agodid you know that this existed and is pretty good and doesn't hog 50GB of swap? https://github.com/iliashad/edit-mind https://github.com/iliashad/edit-mind
- echion 5mo agoThanks for that link -- deserves more attention.
- iliashad 5mo agoAwesome, thank you for mentioning my project. Much appreciated
- dwa3592 5mo agoWhen I saw the original article - i thought this should totally be possible without needing that much ram, so let me just build it. So I started looking at the models I needed for each (whisper, florence etc). Before building, i decided to do one more search on the internet, and I found your project. That's exactly what I was going to build. Good work!!
- iliashad 5mo agoThat's good to know. It was pretty fun to build, thank you!
- mujib77 5mo agoThis is sick. Nice work
- Maya_Andersson 5mo ago[flagged]
- harlanji 5mo agoInteresting. I've been doing similar stuff with my archive on a weak Celeron laptop with 4GB RAM using vanilla ML tech that I'm learning by prompting LLMs (heh). Extract all info from media as sidecar files and all, exploring low power approaches. I can sell this as a service to people who can't even run an LLM, or don't want to cook their hardware. Waitlist open: "Catalog, search, preview, and generate production-ready prompts & scripts from your entire archive — on your existing hardware. Then render in the cloud." https://harlanji.pythonanywhere.com/assetforge/ https://harlanji.pythonanywhere.com/assetforge/
- benbojangles 5mo agoGemma4 because presumably it does image analysis right? -31b It's a dense model -how many tokens/s is it running at -What temps are the M1 max GPU/CPU running at -Is it mlx or gguf -Why 31b and not 26b which is moe and much more efficient on the m1 max at 50tokens/s & low temps. I personally use (MLX) qwen3.6-35b-8bit mostly, but use Gemma-4-26b-4bit for image analysis, its mind blowing how fast it is at identifying the scene in a photograph.
- edg5000 5mo agoLove this article! Had never thought of a use case like this. Had no idea Gemma had a vision encoder. Great use case for local LLM!