Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fabmilo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
fabmilo
2y ago
from my understanding that is what they do, see the paper: > We use a pre-trained GPT-2 (Radford et al., 2019) as the base model for all experiments. I agree the feedback is necessary, and the mechanism simple and cheap, but I don'
32.
▲
by
fabmilo
2y ago
I like the direction of the research of working in latent space but feeding the last layer representation back as a first layer embedding feels sketchy to me. Those layers have different representation space.
33.
▲
by
fabmilo
2y ago
This is interesting as I was evaluating starlark few days ago. The fact that has a customizable implementation in golang, and a python similar syntax makes it an interesting choice for agents generated code.
34.
▲
by
fabmilo
2y ago
Hiring humans to do a consistent job is gonna be a nightmare and a limit on the scalability of the service. How are you defining your service level agreements?
35.
▲
by
fabmilo
2y ago
I really would like to work full time on LLM for code generation. I have many ideas on how to leverage the context length to produce way better output than current models. My current setup is Zed editor + ollama + qwen-2.5-coder on an M3 Ul
36.
▲
by
fabmilo
2y ago
I also thought that they were ending some previous deal and not creating a new one.
37.
▲
by
fabmilo
2y ago
so much pleasantry so much fluff. reduce the noise. get to the point.
38.
▲
by
fabmilo
2y ago
are they hiring?
39.
▲
by
fabmilo
2y ago
Is your website made with repaint? benshumaker.xyz seems could not be indexed by a search engine because everything is draw live. Am I missing something?
40.
▲
by
fabmilo
2y ago
the keyword here is at the end of the article: alarming amount of coffee. There is something neurochemical in highly creative people that needs to be counterbalanced artificially.
41.
▲
by
fabmilo
2y ago
What's the product? a visual code extension with a custom ipython kernel?
42.
▲
by
fabmilo
2y ago
These tools are eye candy and have been around from tensorflow/tensorboard 0.x 10 years ago but never used after just trying them for fun. You need to read the source code no easy way around it.
43.
▲
by
fabmilo
3y ago
I like the decision tree analogy
44.
▲
by
fabmilo
3y ago
I like how it goes from 13.5 GB to 1.76 GB and gets comparable results. Definitely there is some underlying process of why this works so well that we are missing. Are the bits selecting the different subspaces? Can we improve this process b
45.
▲
by
fabmilo
3y ago
I believe this kind of graph exploration is what we need to progress reasoning in AI. Plain LLMS will fail. The link has tons of good references, including the Zobrist hashing https://en.wikipedia.org/wiki/Zobrist_hashi
46.
▲
by
fabmilo
3y ago
They train using Straight Through Estimator but is cited in the previous BitNet paper. What happen to the TrueNorth Chip? I think investing in specialized hardware for AI is a good bet.
47.
▲
by
fabmilo
3y ago
can't you have 2 bits ? first bit for the sign second bit for the 1 0 you can represent -1 +1 +0 -0
48.
▲
by
fabmilo
3y ago
I find this extremely interesting. Do you share the source code of the process? any more references?
49.
▲
by
fabmilo
3y ago
OKRs are not merely tools for creating a set of easily game-able measures for performance review bonuses. Rather, they are a strategic framework designed to establish focus, inspire and build cohesion within teams. The essence of OKRs lies
50.
▲
by
fabmilo
3y ago
Click bait title from someone that doesn't understand the true value of OKRs
51.
▲
by
fabmilo
3y ago
I was about to post that video too. Highly recommended.
52.
▲
by
fabmilo
3y ago
Microsoft just changed the license of phi-2 to MIT!
53.
▲
by
fabmilo
3y ago
Funny I was just looking for something to substitute GPT4V as they are bounding the API usage to few request per day. Sadly this project is built on top of phi-2 that has the non-commercial friendly Microsoft research license.
54.
▲
by
fabmilo
3y ago
This means that we need some new form of preprocessing the data before training LLMs from simple text. Probably just using this compressor and then try to decompress the full text could give some better Supervised Fine Tuned results. Wonder
55.
▲
by
fabmilo
3y ago
This. The best use of the current llms is to create better Datasets.
56.
▲
by
fabmilo
3y ago
I would like to work in this code copilot space, I think will be one of the fastest applications of LLms in the near future. I have been working on a tool to autogenerate docstrings from a python method in google format
57.
▲
by
fabmilo
3y ago
Location: Los Angeles, CA Technologies: Python, PyTorch, AWS Email: mistobaan+yc@gmail.com Looking for training scalable LLM / GPU Deep Learning Kernel optimization using Triton lang/TorchDynamo / model serving using Triton S
58.
▲
by
fabmilo
4y ago
Oh, I like the simplicity of the delivery. I did the same but wanted to publish it on amazon kindle and got lost in the final steps. Maybe I Should just put the website up as you did. Nice graphics, the character consistency is very good, a
59.
▲
by
fabmilo
4y ago
The problem with dating is that we enforce a 1:1 relationship while nature encourages 1:N. Am I the only one seeing this? or we shouldn't talk about it to not risk getting canceled?
60.
▲
by
fabmilo
4y ago
yes, I am going this route and uploading the data to google photos. It will take 7 days to move ~300GB. I discovered that apple uses gcloud underneath and that gcloud/gdrive is the fastest drive service compared to the others. So bough
More ›