Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
SnowflakeOnIce
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
SnowflakeOnIce
3y ago
It's also not clear to me how one would determine that an ML model-generated plan is indeed optimal, or how far from optimal it is. A*-based approaches give you these things.
32.
▲
by
SnowflakeOnIce
3y ago
How would one go about proving that a learned heuristic (something from an AI model) is in fact admissible?
33.
▲
by
SnowflakeOnIce
3y ago
They don't report execution time in the paper. It's likely that the A* implementation would run faster in terms of CPU or wall clock.
34.
▲
by
SnowflakeOnIce
3y ago
Note that the paper doesn't measure execution time of the A* or Transformer-based solutions; it compares length of A* algorithm traces with length of traces generated by a model trained on A* execution traces. I suspect that the actual
35.
▲
by
SnowflakeOnIce
3y ago
I have seen 'file' misclassify many things when running it at large scale (millions of files) from a hodgepodge of sources. Unrelated types getting called 'GPG Private Keys', for example. For textual data types, 'fi
36.
▲
by
SnowflakeOnIce
3y ago
Yes! Sometimes a file has no extension. Other times the extension is a lie. Still other times, you may be dealing with an unnamed bytestring and wish to know what kind of content it is. This last case happens quite a lot in Nosey Parker [1]
37.
▲
by
SnowflakeOnIce
3y ago
As to the your second question: WARC is based on an older format that was started in the 1990s, before SQLite existed.
38.
▲
by
SnowflakeOnIce
3y ago
I think it will stay that way for the foreseeable future (but who can say). Ways to fix the particular hole: (1) disable creating new `code` objects directly from Python. This probably would break lots of things. (2) Add a bytecode verifica
39.
▲
by
SnowflakeOnIce
3y ago
> Conventionally, in software security, Python is considered a memory-safe language. The piece makes the case that Python isn't memory safe when you FFI into a C library. Interesting and largely unknown trivia: it's possible to
40.
▲
by
SnowflakeOnIce
3y ago
I believe MA (which overhauled its noncompete laws a few years back) requires "garden leave" in the case of a noncompete — the employer would have to pay the employee's salary for the period of enforcement.
41.
▲
by
SnowflakeOnIce
3y ago
Surely someone at YouTube could answer this question definitively?
42.
▲
by
SnowflakeOnIce
3y ago
> I’d like to see these many occurrences. Some high-profile cases where credentials were leaked on public GitHub: Uber in 2014 and 2021 [1, 2] and Twitch in 2021. > Search is still possible using google and other methods Yes, you can
43.
▲
by
SnowflakeOnIce
3y ago
Yes, you are correct. But requiring that one be logged in to use the search functionality makes a stronger audit trail possible: looking at the search logs would indicate who has been hunting for secrets!
44.
▲
by
SnowflakeOnIce
3y ago
Another factor: anonymous faceted regex search across a huge volume of code allows bad actors to find hardcoded credentials and gain access to additional systems, without a good audit trail. But yes, there are multiple good explanations for
45.
▲
by
SnowflakeOnIce
3y ago
Aside from performance concerns, there have been many occurrences of bad actors using code search to find hardcoded credentials, and then using that to gain unintended access to additional systems. I suspect the change has more to do with s
46.
▲
by
SnowflakeOnIce
3y ago
The article is wrong. It's an hour northwest of Boston.
47.
▲
by
SnowflakeOnIce
3y ago
This is not quite a command palette, but in macOS apps, you can press command-shift-/ to search the names of menu commands. I use that frequently in apps that I don't use regularly (and hence don't remember where all the men
48.
▲
by
SnowflakeOnIce
3y ago
Depending on many factors (like details of the patterns used and the input), some regex engines (like Hyperscan) can match tens of gigabytes per second per core. Shockingly fast!
49.
▲
by
SnowflakeOnIce
4y ago
'pip freeze' will generate the requirements.txt for you, including all those transitive dependencies. It's still not great though, since that only pins version numbers, and not hashes. You probably don't want to manually
50.
▲
by
SnowflakeOnIce
4y ago
The C API for Python changes between major Python versions. This means that native-code extensions (like torch) need to be build specifically for each version of Python they want to support.
51.
▲
by
SnowflakeOnIce
4y ago
It's more an issue of torch not yet providing prebuilt binary wheels for Python 3.11. You could probably get it working, but it would involve building torch from source, which can be rather more involved than 'pip install'. I
52.
▲
by
SnowflakeOnIce
4y ago
Yes and no. On the one hand, Nosey Parker is effectively a special-purpose `grep` with a bunch of security-relevant patterns built-in, including one for PEM-encoded keys: < https://github.com/praetorian-inc/noseyparke
53.
▲
by
SnowflakeOnIce
4y ago
That's an interesting use case! The development so far has focused on security-related uses, especially finding hardcoded credentials in source code and log files. For the serial number and license spelunking you describe, Nosey Parker
54.
▲
Show HN: Nosey Parker, a fast and low-noise secrets detector for textual data
(github.com)
65 points
by
SnowflakeOnIce
4y ago
|
6 comments
55.
▲
by
SnowflakeOnIce
4y ago
This reasoning works okay until you have to deal with machine-generated inputs, at which point your app freezes when it gets sent down a quadratic or worse complexity path.
56.
▲
by
SnowflakeOnIce
4y ago
I understand your point and perhaps I'm in fierce agreement. OTOH, things like parser combinators are much more ergonomic in Haskell than Rust.
57.
▲
by
SnowflakeOnIce
4y ago
My experience with a modest machine and a ~100GB dataset was that DuckDB was significantly easier to use and much faster (20x) than Ray Data. Have not compared directly with Spark. There was no cluster to set up or administer using DuckDB.
58.
▲
by
SnowflakeOnIce
4y ago
DuckDB is terrific. I'm bullish on its potential for simplifying many big data pipelines. Particularly, it's plausible that DuckDB + Parquet could be used on a large SMP machine (32+ cores and 128GB+ memory) to deal with data mung
59.
▲
by
SnowflakeOnIce
4y ago
What vector database do you use? Did you run into any scalability challenges there?
60.
▲
by
SnowflakeOnIce
4y ago
Very interesting! > Strangely, these bugs were found by the CI of ClickHouse, and not by any of the hundreds of other products using these libraries. A lot of software projects that have "OSS Fuzz integration" have very naive o
More ›