Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aleksiy123
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
aleksiy123
13d ago
As a sidenote from the slop language, which is awful. Feel like latest models have really got hung up on “facts” and “evidence” in a really unnatural way. And especially Claude seems to have a lot of trouble separating it’s internal thought
2.
▲
by
aleksiy123
15d ago
Tyty I was thinking about something similar
3.
▲
by
aleksiy123
17d ago
You can also get LLM to optimize rules for a rules engine iteratively against some dataset. It’s sort of like memoizing or distilling the knowledge. Works really well for certain type of problems.
4.
▲
by
aleksiy123
1mo ago
The font, the colour, and the dot on the tags all scream Claude slop ui.
5.
▲
by
aleksiy123
1mo ago
PlayStation store not allowing returns for games that have been downloaded is the most bs policy. In a world where steam allows instant returns with under 2 hours play time. I have a few games now on ps where I barely played or found out th
6.
▲
by
aleksiy123
1mo ago
You can do much more with feature flags like ab testing/experiments, integrations with analytics. There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer
7.
▲
by
aleksiy123
1mo ago
Managing complexity while scaling is a concept just keeps reappearing. Funnily enough it’s extremely relevant to agents/subagents.
8.
▲
by
aleksiy123
1mo ago
Huh this is cool. Was thinking about earlier that it must be a lot easier to make motion capture games now. Are there any new dance style games? As an aside wonder if there is something to adding game/gambling mechanics to exercise.
9.
▲
by
aleksiy123
1mo ago
Just means it can be powerful or useful technique that can feel “magical” when employed. I wouldn’t overthink this.
10.
▲
by
aleksiy123
1mo ago
“Extreme defensiveness” is a good characterization. I wonder if it’s an artifact of OpenAI’s values or rl training approach. Also, it prob does make it perform better just not more efficient. Great for the OpenAI employee working on securit
11.
▲
by
aleksiy123
1mo ago
As a follow up. I feel like codex/sol is better at well scoped hard technical problem. Where it can sort of run this brute force analytical loop. Like doing performance optimization or other search type problems. I think the math proof
12.
▲
by
aleksiy123
1mo ago
Agree with most of these. One thing I don’t love about codex/sol is I find it tends to overengineer and be overly cautious. I was using it to do create some scraping + data processing. It went kind of crazy on the provenance, need at l
13.
▲
by
aleksiy123
1mo ago
Hooks is the way. Intermittent nudges
14.
▲
by
aleksiy123
2mo ago
I use a hook. Checks for banned words and patterns. Injects a reminder to use ASD-STE100 Simplified Technical English which I picked up from a suggestion in another thread. Honestly it’s working pretty well. Except for I need to check how o
15.
▲
by
aleksiy123
2mo ago
¯\_(ツ)_/¯ You can believe whatever you want to believe. Or maybe you live in a magical world where all your tests are pure and side effects nor IO exist. In the other thread you think running tests in parallel doesn't count becaus
16.
▲
by
aleksiy123
2mo ago
You’ve never ever had to be mindful about isolation of tests? Also, Unit tests aren’t the only tests that exist. You say assured? You think it’s impossible to write 2 tests that interact each other? All you gotta do is google “test isolatio
17.
▲
by
aleksiy123
2mo ago
This
18.
▲
by
aleksiy123
2mo ago
All I mean that there is no way for a framework to prevent your tests from interfering with each other, or to solve isolation for you. It’s the implementation of each test that is responsible for its isolation. Every time you write or make
19.
▲
by
aleksiy123
2mo ago
I guess it all depends, isolation isn’t binary. you can share some things and be isolated across other dimensions. You could share a Postgres dB connection but just isolate the data logically. You can isolate or share across individual test
20.
▲
by
aleksiy123
2mo ago
I don’t think that’s right? The isolation comes from the test implementation not the framework. There isn’t any framework out there that can guarantee/give you isolation. If I create a new in mem dB in the test there’s nothing stopping
21.
▲
by
aleksiy123
2mo ago
Is this only for realtime tts use cases? Wondering if you also support some non realtime models.
22.
▲
by
aleksiy123
2mo ago
Curious if you can prompt Claude to sue some scrambling scheme and then unscramble to defeat this. E.g. prompt Claude to write all sentence in reverse, or swap every 2 words etc. Then use a script to put reorder in the right ordering?
23.
▲
by
aleksiy123
2mo ago
Jeopardy clustering Pretty cool technique honestly. You could do it the other way as well right? If you had a list of categories you have the model to generate a sample query and then do embedding on that?
24.
▲
by
aleksiy123
2mo ago
Is it possible to have some kind of script to keep your cache warm, or auto compact or something. I sometimes just leave some goals or something running before I go to bed or out and I don’t want to pay the cache text when I come back.
25.
▲
by
aleksiy123
2mo ago
but is it even possibly to selectively apply it? I guess some kind of tag or indicator token that its code or not code?
26.
▲
by
aleksiy123
2mo ago
In curious how does this work with tool calls or CLI scripts etc? like if you have a long cli command or something will it still try to watermark it ? Is there some way you can know which tokens are required to be syntactically correct vs n
27.
▲
by
aleksiy123
2mo ago
I can make a decision to go to gym tomorrow. My process/playbook for determining if I go to the gym is I go if I didn’t go the day before. This is both a premade decision and a process that requires discipline to execute on. So why isn
28.
▲
by
aleksiy123
2mo ago
This is a good one. Sometimes it doesn’t even matter if the made decision is suboptimal. It’s better than no decision. Opinionated linters is an example. Sometimes you just gotta turn your brain off and move forward.
29.
▲
by
aleksiy123
2mo ago
I mean people don’t usually even do 8 hours of deep focus. Also some things just have wall time. Sometimes it may take 3-4 hours to roll things out and test. Babysit some dashboards, cicd. Working “12” hours may mean you get 3-4 attempts&#x
30.
▲
by
aleksiy123
2mo ago
Some of these are really lacking in gravitas.
More ›