2 ms·
While I agree that the expectations for AI are too high, I think the use case described is almost achievable with today's tech. It's just language processing to
by huntertwo 5y ago
While I agree that the expectations for AI are too high, I think the use case described is almost achievable with today's tech. It's just language processing to process the request (User wants to see dog around this time period in this location) and finding relevant images (find pictures with dog with exif data that matches time period and location).
- gwern 5y agoYou could do most of that with some creative use of CLIP. 'Show more of the mountain' is trivial out of the box, and the 'kids aged 5-10' just needs a bit of logic: take a few photos of each kid at various ages 5-10 (perhaps already specified by the user?), embed, average along with the text prompt 'kids aged 5-10', and use those as the targets. 'Around that time' and 'in and around that house' can be done by metadata dates + GPS + image embedding of the 'house' (or just the text embedding of the word 'house'!) to filter photos by date+location, and then similarity to 'house'. And so on. The question there is how to be able to create those queries automatically in response to transcribed user voice commands. But hey, given Codex/Copilot, who can doubt that programming such queries is out of reach of a contemporary NN either?