Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eliomattia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Magic LTM-1, a coding assistant LLM with a 5M token context window
(magic.dev)
5 points
by
eliomattia
3y ago
|
0 comments
2.
▲
Google researchers: non-editable LLMs can be modified by chatting, no finetuning
(arxiv.org)
4 points
by
eliomattia
3y ago
|
0 comments
3.
▲
by
eliomattia
3y ago
The article can be summarized as: context richer prompt history yields answers that are better aligned with expectations. > the practical constraint of a finite context window coupled with meta-in-context learning’s rapid prompt length i
4.
▲
by
eliomattia
3y ago
By human feedback in T, I meant indeed RLHF. By chat histories H in T, I meant a later selection of user feedback. While plug-ins and added context can be visualized as g(f(H)), fine tuning could be thought of as g(f)(H). There are humans i
5.
▲
by
eliomattia
3y ago
> This is true for almost everything politicians promote as a "solution" :-) In their defense, this time they may not even be aware. > So for most people, they actually see g which is something like g(f(H))) It is interestin
6.
▲
by
eliomattia
3y ago
There are two divergent analogies to analyze. Compared to programming software, LLMs are experientially closer to the malleability of interacting with humans on the one hand, while functionally closer with hardware entailment-wise, on the o
7.
▲
by
eliomattia
3y ago
There are assumptions here that are intimately related to the meta questions mentioned. > Prompt injection or prompt attacks are well known and likely impossible to guard against. They are impossible to guard against under the assumption
8.
▲
White House is working with hackers to ‘jailbreak’ ChatGPT’s safeguards
(fortune.com)
1 points
by
eliomattia
3y ago
|
13 comments
9.
▲
by
eliomattia
3y ago
As a programmer, I find it fascinating to build things from the ground up, with the inner workings either in full display or readily accessible for editing. With AI, the need to beg it to please behave, with a long list of things to do and
10.
▲
by
eliomattia
3y ago
Really interesting, also diff-based, and 3.5 years in development. On the homepage I read "An in-memory, distributed, and open-source document graph database". Do you know whether the whole database, including all documents, needs
11.
▲
by
eliomattia
3y ago
I just found this blog post. It seems Palantir Foundry, which does not come up often when researching git for data tools, includes a version control system for datasets that stores diffs in their own cloud-based filesystem. According to the
12.
▲
Palantir Foundry's dataset version control, a diff-based Git for data
(blog.palantir.com)
3 points
by
eliomattia
3y ago
|
4 comments
13.
▲
Ask HN: How to name nested not well-defined sets and nested cyclic metagraphs?
2 points
by
eliomattia
3y ago
|
0 comments
14.
▲
by
eliomattia
3y ago
Fully agree. Compression in many cases removes the ability to diff easily, however. In a large dataset where, in terms of size, 1% of the original data undergoes changes, or new data the size of 1% of the original dataset is added, I think
15.
▲
by
eliomattia
3y ago
*I have just (only now) read the second paragraph in your message. Not sure if that came across correctly, that first sentence was too compressed.
16.
▲
by
eliomattia
3y ago
The two are not mutually exclusive, in principle. Depending on workflows, sizes, and change frequencies, each has advantages. Sparse checkouts are useful with small files that specific individuals can focus upon, among other things. This is
17.
▲
by
eliomattia
3y ago
Just read the second paragraph. Currently expanding merge resolution assistance to deal with the general merge conflict case, as well as implementing revert and cherry-pick assistance. Unsure if that is what you were wondering? You probably
18.
▲
by
eliomattia
3y ago
Sorry I should have been more specific, I meant block deduplication, or any form of deduplication at a level lower than the entire file. File deduplication can only get you so far, depending on the use case. XetHub does block deduplication,
19.
▲
by
eliomattia
3y ago
That is really interesting and begs the question of how frequently you have changes in your data that lead to new commits. I am assuming here that you don't dedupe anything, that is, you throw the entire files into Azure with each vers
20.
▲
by
eliomattia
3y ago
Git does do compression on repos, but the fact that versioning repositories with (huge) data is still an open problem suggests that it is not the kind that fixes it. I might be mistaken, are you aware of any interesting compression methods
21.
▲
by
eliomattia
3y ago
Which repo size after the filters do you work with on your machine and how many GBs do you have in Git LFS, that is, in the cloud? I hear people complain about costs, but it depends upon scale and change frequency, which can increase total
22.
▲
by
eliomattia
3y ago
It's like GVFS, but for pieces of a file at a time as well: rows, columns, or cells. A snapshot is recreated by putting those pieces together. If you have ten million rows in one file and only add a thousand rows daily, each commit wil
23.
▲
Committing changes to a 130GB Git repository without full checkouts [video]
(youtube.com)
20 points
by
eliomattia
3y ago
|
16 comments
24.
▲
by
eliomattia
3y ago
Full checkouts of large data repositories are problematic. In the video I present a workflow that does not require full checkouts of the datasets and still allows to commit diff-based changes in Git. This naturally applies to new data, and
25.
▲
Show HN: Git for datasets and config table versioning, that commits diffs
(dropbox.com)
1 points
by
eliomattia
3y ago
|
0 comments
26.
▲
by
eliomattia
3y ago
I am working on a solution on top of Git, but storing diffs only, can integrate with MySQL and S3, can create version snapshots. You would have the videos in a bucket and the version-controlled history with links to the videos in another, a
27.
▲
by
eliomattia
3y ago
Important points. I aim for version control for data repositories with HDD efficiency, visualization of diffs for collaboration, and API accessibility of individual datasets from multiple identifiable versions from git. Datasets can then be
28.
▲
by
eliomattia
3y ago
dolt came up when searching git for data, it seems great, though I have never used it. I know it works on prolly trees rather than on top of git. I am really curious to learn about that choice, why exactly not on git? How can you offer data
29.
▲
by
eliomattia
3y ago
I have built version control for data, on top of git itself, that can commit and push incremental diffs. By tagging in git, a version snapshot can be created. S3 can be configured (a) for heavy files and diffs referenced by pointer objects
30.
▲
Git for data that commits incremental diffs to Git itself
(youtube.com)
2 points
by
eliomattia
3y ago
|
5 comments
More ›