4 ms·
This feels like the peak of resume driven development. The maker of this has taken a deterministic problem (substring matching transaction descriptions) that co
by Fiveplus 9mo ago
This feels like the peak of resume driven development. The maker of this has taken a deterministic problem (substring matching transaction descriptions) that could be solved with a 50-line Python script or a standard .rules file and injected a non-deterministic, token-burning probability engine into the middle of it.
I'll stick to hledger and a regex file. At least I know my grocery budget won't hallucinate into "Consulting Expenses" because the temperature was set too high.
- ezst 9mo agoDude, common, we won't be reaching "AGI" with that attitude. /s
- deleted 9mo ago[deleted]
- nerdzoid 9mo agoI think you don't use UPI or you would have understood the painpoint categorizing even with ai would be difficult.
- minne 9mo agoThe maker built NuGet. He don't need a resume.
- pansa2 9mo agoBut can he invert a binary tree?
- junto 9mo agoAgreed but not just NuGet. David also was a key driver for SignalR and now Aspire (which is genuinely the most awesome tool I’ve seen for a while). He’s also extremely humble and doesn’t need to try to impress anyone. I think it’s clear that he’s just doing this tool for fun and chose to share it. People shouldn’t mix up their anti-Microsoft autoeroticism and a person that happens to work for them.
- throwaw12 9mo agoThat dude is a Distinguished Engineer at Microsoft, doesn't need your "resume driven" label, his resume is good enough already. Why don't you accept it as, dude is experimenting and learning new tool, how cool is that, if this is possible, what else can I build with these tools?
- supriyo-biswas 9mo agoThe pressure at those levels is even higher, as it is an unsaid expectation of sorts that LLMs represent the cutting edge of technology, so principals/DEs must use it to show that they're on the top of the game.
- throwaw12 9mo agofair, but it doesn't mean some of them are genuinely experimenting and figuring out interesting ways to use LLMs, some examples I personally love and admire * simonw - Simon Willison, he could just continue building datasette or help Django, but he started exploring LLMs * Armin Ronacher * Steve Yegge and many more
- andy99 9mo agoNo idea if this is true but very sad if it is. This is a great argument for the concept of tenure, so experts can work on what they as experts deem important instead of being subject to the whims of leadership. I, probably naively pictured Distinguished Engineer to be closer to that, but maybe not.
- yolo3000 9mo agoIt's in the career framework of most big techs to use AI this year, so everyone is doing it to hold on to their bonuses.
- RevEng 9mo agoSadly, yes, it's true. New AI projects are getting funded and existing non-AI projects are getting mothballed. It's very disruptive and yet another sign of the hype being a bubble. Companies are pivoting entirely to it and neglecting their core competencies.
- makach 9mo agoThis is an actual hard problem he is trying to fix. 50-line python? Pfff..! My current personal 400-line rule script begs to differ not to mention the PAIN of continuously maintaining it. I was looking into using AI to solve the same problem but now I can just plug and play.
- deleted 9mo ago[deleted]
- senko 9mo ago> a deterministic problem (substring matching transaction descriptions) that could be solved with a 50-line Python script I don't know about your bank transactions, but at least in my case the descriptions are highly irregular and basically need a hardcode for each and every damn pos location across each of the store location across each vendor. I attempted that (with a Python script), gave up, and built myself a receipt tracker (photo + gemini based ocr) instead, which was way easier and is more reliable, even though - oh the horror! - it's using AI.
- noufalibrahim 9mo agoIsn't this suitable for a Bayesian classifier? Label some data (manual + automation using substrings) and use that to train a classifier and then it should be able to predict things for you fairly well. There's a new feeling that I experience when using LLMs to do work. It's that every run, every commit has a financial cost (tokens). Claude code can write a nice commit message for me but it will cost me money. Alternatively, I can try to understand and write the message myself. Perhaps the middle ground is to have the LLM write the classifier and then just use that on the exported bank statements.
- yunohn 9mo agoIMHO the better middle ground is to use a nice (potentially fine tuned) small model locally, of which there are many now thanks to Chinese AI firms.
- ithkuil 9mo agoAn expensive model can generate the training dataset
- ozim 9mo agoGetting LLM to write the classifier should be the way to go. That’s what I mostly do, I give it some examples ask to write code to handle stuff. I don’t just dump data into LLM and ask for results, mostly because I don’t want to share all the data and I make up examples. But it also is much cheaper as once I have code to process data I don’t have to pay for processing besides what it costs to run that code on my machine.
- andy99 9mo agoAnd this feels like peak HN “why not just use a regex”. This is a hard, for all intents and purposes non deterministic problem. Now if you’ll excuse me I have to draft another post admonishing the use of curl | sh
- hrimfaxi 9mo agoIt seems like using a model to create regexes that match your transactions might be worthwhile.
- ManuelKiessling 9mo agoYeah, a pattern like „do the heavy lifting with cheap regexes, and every 100 line items, do one expensive LLM run comparing inputs, outputs, and existing regexes to fine-tune the regexes“.
- dizhn 9mo ago> The maker of this has taken a deterministic problem (substring matching transaction descriptions) that could be solved with a 50-line Python script I had a coding agent write this for me last week. :D It takes an excel export of my transactions that I have to download obviously since no bank is giving out API access to their customers. It uses some python pandas and excel stuff and streamlit to classify the transactions and "tally" results and show on the screen as color coded tabular data. (Streamlit seems really nice but super super limited in what it can do.) It also creates an excel file (with same color coding) so I can actually open it up and check if necessary. This excel file has dropdowns to reclassify rows. The final excel file also has formulas in place to update live. Code can also compare its own programmatic calculations with the result from formulas from excel. Why not? My little coding sweatshop never complains. (All with free models and clis by the way. I haven't had a reason to try Claude yet.)
- xtiansimon 9mo ago> "...no bank is giving out API access to their customers..." I think Citi bank has an API (https://developer.citi.com/ https://developer.citi.com/). Not that it's public for account holders, but third-parties can build on that. I'm looking at Plaid.com One thing about Plaid--I've not been happy when encountering Plaid in the wild. For example, when a vendor directs me to use Plaid to validate a new bank account integration. I'd much rather wait a few days and use the 0.01 deposit route. Like using a katana for shaving. But signing up to use Plaid for bank transaction ingestion via API is a whole different matter.
- zwnow 9mo agoCan't wait for people building agent based grep or whatever else solved issues there are. AI people really need to touch grass.
- dimitri-vs 9mo agoIn theory, yes. In practice the shit data you are working with (descriptions that are one or two words or the same word with ref id) really benefit from a) an agent that understands who you are and are likely spending money on b) has access to tool calls to dig deeper into what `01-03 PAYPAL TX REFL6RHB6O` actually is by cross referencing an export of PayPal transactions. I think the smarter play is having an agent take the first crack at it, and build up a high confidence regex rule set. And then from there handle things that don't match and do periodic spot checks to maintain the rule set.
- xtiansimon 9mo ago> "...then from there handle things that don't match..." Curious, what's the inputs for an agent when handling your dataset? What can you feed the agent so it can later "learn" from your _manual labeling_?
- init 9mo agoI've built and worked on this exact problem before at bigtech, startup and personal projects. Regex works well if you have a very limited set of sender and recipient accounts that don't change often Bayesian or DNN classifiers work well when you have labeled data. LLMs work well when you have a lot of data from lots of accounts. You can even combine these approaches for higher accuracy
- conradev 9mo agoI imagine a coding agent would be great at editing your regex file to maximize coverage. Just like manually editing Sieve/Gmail filters: I want full determinism, but managing all of that determinism can be annoying…
- qaboutthat 9mo agoCharitably, this is a very naive take for unstructured bank / credit card transactions. Even if you use a paid service for elaboration you will not write a 50-line, or even 500-line, list of declarative rules to solve this problem.
- froggertoaster 9mo agoLmao I'm no David Fowler fan (leftist blowhard), but he's one of the most talented and successful engineers at Microsoft. I don't think he needs to build a resume.