Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tsvoboda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
tsvoboda
8mo ago
Would love to hear how you're handling recovery for long-running training jobs today, as well as what failure modes are most common/annoying for you.
2.
▲
Show HN: Autonomous recovery for distributed training jobs
(docs.tensorpool.dev)
12 points
by
tsvoboda
8mo ago
|
3 comments
3.
▲
by
tsvoboda
1y ago
looks pretty cool! How would you integrate this into production agent stacks like langchain, autogpt, even closed loop robotics?
4.
▲
by
tsvoboda
2y ago
this is dope, i hate maintaining custom integration code
5.
▲
by
tsvoboda
2y ago
pretty sick stuff guys, excited to see what you accomplish