4 ms·
So there are models for Pose estimation like [PoseNet] or Qualcomm's [MediaPipe] that give you a graph of the body position. This metadata can be converted to a
by miguelff 2y ago
So there are models for Pose estimation like [PoseNet] or Qualcomm's [MediaPipe] that give you a graph of the body position. This metadata can be converted to a string representation and be fed into a generative model. If I record myself in different exercises, extract the poses, and label them with the appropriate steps, I might be able to perform few-shot inference or RAG over a set of labeled positions, such as the model can infer new body positions from exercise steps.
If that'd possible, then it is fairly easy:tm: to generate illustrations (e.g. like stickman figures. See [Kawamoto] for an example)
Of course, none of this is necessary if I can record (or let users record) and upload videos of each exercise, having enough of them to not feel they are repetitive.
[PoseNet](https://github.com/tensorflow/tfjs-models/tree/master/posenet)
[MediaPipe](https://huggingface.co/qualcomm/MediaPipe-Pose-Estimation)
[Kawamoto](https://github.com/kenkawakenkenke/stickfigure-recorder)