3 ms·
It cannot write some kind of short texts either. The experiment was: given existing Experience section of Senior Software Engineer CV (manually and carefully w
by ink-splatters 2mo ago
It cannot write some kind of short texts either.
The experiment was: given existing Experience section of Senior Software Engineer CV (manually and carefully written), write 2-4 lines of high pitch About [me].
gpt-sol-xhigh just could not do it, making a complete AI slop mess.
It became substantially better when I asked to get a sample of real Senior/Principal engineers CVs, that it would be able to connect with meaningful online presence/contributions of their authors, and draw ideas from that.
Nevertheless, the remaining issues were critical, with their classes spanning:
- word for word repetition;
- tautology (phrases mapping to the same semantic entity);
- category mistakes (combining apples with oranges);
- faulty composition of generalised and concrete terms.
Before throwing it away I decided to give it a try and asked for strict prompt following, setting low logical errors threshold, and eventually providing a formal proof that it complied.
It took spaCy, doing NER and dependency parsing; then I suggested adding stanza for constituency parsing.
After slapping together the artefacts of analysis and thinking a bit, it produced great phrase (to my taste).
-
Then when I asked to apply the framework and improve bullet points in some experience block, it did 2 of 4 well, then miserably broke down; I’m not sure if it’s harness issue (codex) or fundamental model restriction.
-
So no, without substantial investment in steering, SOTA AI doesn’t perform even remotely close to a human, in complex reasoning.
- swat535 2mo agoIt's funny, because right now the top news on HN is about OpenAI's "Ten advances in mathematics and theoretical computer science". So I suppose which is it? Is AI going to replace us all, or it's not even capable of writing basic code? I've never seen a topic so divisive where people's experiences differ in such magnitudes. I suppose the biggest thing drawing this distinction is that on one side, you have AI releasing features which is all the business cares about. On the other side you have it generating slop, which is all engineering ever cares about. I think in this battle, executives are going to eventually win. They had been looking for a way to crush the engineering department's value for years and AI seems to be their golden ticket. The pesky dev cost center can finally shrink and they don't have to bother paying them six figures thanks to LLMs..