4 ms·
Show HN: Littlebird – Screenreading is the missing link in AI
- deleted 6mo ago[deleted]
- wolfduck 6mo agoLooks cool! Reminds me of Rewind but better. (edit: for completion)
- grena1re 6mo agoHi all, I'm Alex, one of the founders at Littlebird. AMA!
- winterbloom 6mo agoDoes your company actually look at resumes? you post in HN but nothing ever goes through. I have a feeling that your recruitment team is not doing their job properly (or being super selective)
- winterbloom 6mo agoI guess it is AMA, anything except about your current hiring pipeline lol
- deleted 6mo ago[deleted]
- Remi_Etien 6mo ago[flagged]
- grena1re 6mo agoLittlebird is a desktop app that remembers everything you’ve been working on. Meetings, messages, docs, browsing, etc. It helps you stay focused, prioritize, recall, and move projects forward. Unlike any product on the market today, Littlebird uses screenreading to understand all the text on screen, for all applications, without any cumbersome setup. It understands who said what, when, and keeps track of your projects in great detail. It uses that context to build a rich understanding of your life: who matters to you, what you're working on, and what you care about this week and this year. It extends your working memory and your capacity to think and create. You control what Littlebird sees, what it remembers, and what it forgets. We designed Littlebird to be private, secure, and user-controlled by default.
- anonhn58 6mo agoHi folks, Tushar here from the engineering team at Littlebird. AMA!
- johnsmith1840 6mo agoEver consider the enclave route for this kind of work?
- r3gal08 6mo agoHow are you handling the data extraction? Is it a multimodal VLM (OCR+LLM) or a standard OCR engine feeding a separate LLM? I’ve been hitting a wall trying to understand how this viable. The compute overhead for real-time analysis at scale seems massive without a serious backend. How are you managing the frequency?
- anonhn58 6mo agohi, while vision is going to be a part, it's hard to scale both on server and the client(it's resource intensive and the battery will drain faster on client). we hook deeper into the OS layer with accessibility, Apple Script and other ways to get raw text. This also lets us create a privacy friendly app with granular data controls for the user. > compute overhead for real-time analysis at scale seems massive without a serious backend. you're still right about this part though, and we do have a serious backend.
- ycyash 6mo agotried this last week. typed three words, got back a full breakdown that pulled from three different tools i'd had open. didn't brief it once. Good tool. Any plans for Windows ? It's my main workstation.
- antonholub 6mo agohey! yes, version for Windows is in active development and if everything is fine it will got to beta soon.
- Achiyacohen 6mo ago[dead]
- rsingel 6mo agoIf you thought Slack logs were damning in discovery, wait til someone suing or prosecuting you figures out that everything you typed and looked at, etc., is in the cloud
- edwardsrobbie 6mo ago[flagged]
- divmain 6mo agoIs there any chance you might support a local-first version of this in the future? I've been interested in apps like this and Littlebird in particular seems very attractive. But I'm loathe to essentially send screenshots/summaries/etc of all my activity to a cloud solution, regardless of any claims you make about encryption. Any mistake you make could be catastrophic for me, which thoroughly dominates any upside to using your product. It's a non-starter.
- grena1re 6mo agoWe will for sure, but the issue is that without local LLMs, there's no way to offer a truly fully local version. And the local LLMs are dumb. So basically, you would still need to trust the LLM providers. Totally understand that this is a deal breaker for some people, but for many users, the theoretical risk is worth it. We do regular security audits, encrypt in transit and at rest, pen tests, etc.
- throwaway-blaze 6mo agoUm, dismissing the tech as "the local LLMs are dumb" seems shortsighted. I can run some pretty impressive models on my local Mac, but it has >64gb of ram and an M3 Max. Given the privacy benefit I wouldn't dismiss them so fast. I'd suggest picking one or two that your prompts will work well with and treating it as "we let you run with local models too, if you have a computer capable of that." This will (a) quiet the people who complain about everything and (b) get more people to try the cloud model knowing they could move to a local model for real usage.
- grena1re 6mo agoI'm not dismissing them. I'm saying they're not there yet. As a startup, we have to prioritize. We can't do everything simultaneously, and it would be a substantial engineering effort to have a dual architecture as well as potentially more security holes. And the amount of people that want to run local LLMs is very small. I use local LLMs when I'm on flights, and that is my personal assessment. They are all benchmark-maxed and incapable of reliable tool calling or consistency over meaningfully long conversations.
- Paulo75 6mo agoScreenreading is a smart way to solve the integration problem. Every other tool in this space makes you connect each app one by one and you're always waiting for them to support your workflow. This just watches what you watch. Feels obvious in hindsight - cool stuff
- steve-atx-7600 6mo agoI wish Claude cowork could get better at this. I often end up with Claude performing ad hoc tasks involving multiple windows and it’s so slow. It stops and screenshots and thinks and screenshots to confirm after entering one Google sheet cell of data for example. I’m sure it’ll get better over time.
- reverius42 6mo agoIsn't this a lot like Microsoft's Windows 11 Recall feature that they got a lot of flak for?
- munio 6mo agoThe screenreading approach is genuinely clever as a distribution strategy — no integrations to maintain, no OAuth flows to break, works on day one with everything. The hard part isn't the tech though, it's the trust problem. Rewind, Limitless, and now this all hit the same wall: the people most likely to benefit from this (busy professionals with complex workflows) are exactly the people most exposed if something goes wrong. Until there's a credible local-first path, the TAM is going to stay small.
- alexovch 6mo agoAgree with this direction. Most tools assume structured inputs, but real workflows are messy and visual. Feels like screen understanding is still very underexplored compared to text.
- mcheemaa 6mo ago[dead]