Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
WiSaGaN
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
WiSaGaN
1y ago
This actually makes sense because in the meantime Meta is ditching the open-source (open-weights) direction. Before the national security narrative took over, the main argument was about "safe" AI, where releasing models as open w
32.
▲
by
WiSaGaN
1y ago
I am wondering how much of this can be mitigated by carefully designing feature flags, and make default feature set small.
33.
▲
by
WiSaGaN
1y ago
Can you name one example of a consumer product that China initially sold affordably to gain market share and then later raised prices?
34.
▲
by
WiSaGaN
1y ago
Deepseek and Alibaba just published their froniter models in open weights weeks ago. And they happen to be the leading open weights models in the world. What are you talking about?
35.
▲
by
WiSaGaN
1y ago
Deepseek published thinking trace before OpenAI did, not after.
36.
▲
by
WiSaGaN
1y ago
I don't think mentioning Rust on an article specifically talking about a memory safety bug count as "constant". This is Rust's core strength.
37.
▲
by
WiSaGaN
1y ago
Interesting that the output price per 1M tokens is $0.6 for non-reasoning, but $3.5 for reasoning. This seems to defy common assumption of how reasoning models work, and you tweak the <think> token probability to control how much thin
38.
▲
by
WiSaGaN
1y ago
This common feature requires the user of the API to implement the tool, in this case, the user is responsible to run the code the API outputs. The post you replied suggests that Gemini will run the code for the user behind the API call.
39.
▲
by
WiSaGaN
1y ago
Also the "project" feature in claude improves experience significantly for coder, where you can customize your workflow. Would be great if gemini has this feature.
40.
▲
by
WiSaGaN
1y ago
Deepseek doesn't need hype to survive. They are bankrolled by their now billionaire founder.
41.
▲
by
WiSaGaN
2y ago
This may make sense if there is a centralized force to dictate how much these Chinese foundational model companies charge for their models. I know in the west people just blanketly believes that the state controls everything in China. Howev
42.
▲
by
WiSaGaN
2y ago
I have always suspected that the o1-Pro is some kind of workflow on the o1 model. Is it possible that it dispatches to say 8 instances of o1 then do some type of aggregation over the results?
43.
▲
by
WiSaGaN
2y ago
What user experience are you talking about? Chatbot? Or software in general? Cause Tiktok beats Facebook out of water. Chatbot for English communities sure, I also prefer Claude over Deepseek in terms of project support and UI. But this is
44.
▲
by
WiSaGaN
2y ago
Yes, it is fascinating that humans can have such seemingly fundamental differences in how they function 'under the hood.' I also have a friend who is highly intelligent—they earned a STEM PhD from one of the best universities in t
45.
▲
by
WiSaGaN
2y ago
Before Deepseek, Meta open-sourced a good LLM. At the time, the narrative pushed by OpenAI and Anthropic was centered on 'safety.' Now, with the emergence of Deepseek, OpenAI and Anthropic have pivoted to a national security narra
46.
▲
by
WiSaGaN
2y ago
I usually add "upgrade": `uv tool install --upgrade --python python3.12 aider-chat`. So that it will upgrade the version to the latest if the current one is not already.
47.
▲
by
WiSaGaN
2y ago
I think it will be more akin to o1-mini/o3-mini instead of r1. It is a very focused reasoning model good at math and code, but probably would not be better than r1 at things like general world knowledge or others.
48.
▲
by
WiSaGaN
2y ago
Glad to hear they will revive the NBA part. Was using the model extensively. It was very informative.
49.
▲
by
WiSaGaN
2y ago
Does this require nightly? If so, #[warn(clippy::wildcard_enum_match_arm)] will do the samething but no need for nightly, and from clippy instead of rustc natively.
50.
▲
by
WiSaGaN
2y ago
Indeed, the blog mentioned in the other comment showed part of 3FS code was completed at least since 2019, when this was still a project of the quant funds. In HFT, you tend to dogfood a lot of the things to achieve low latency, high perfor
51.
▲
by
WiSaGaN
2y ago
I think these kind of open-source is really showing their objective of achieving efficiency in the industry. The reason is this kind of software benefits a lot to the big guys serving the model (competitors to Deekseek themselves if they ar
52.
▲
by
WiSaGaN
2y ago
My personal experience is that R1 is smarter than 3.5 sonnet, but 3.5 sonnet is a better coder. Thus it may be better to let R1 to tackle the problem, but let 3.5 sonnet to implement the solution.
53.
▲
by
WiSaGaN
2y ago
It would be hilarious if this scenario played out. OpenAI starts as a nonprofit, aiming to benefit all humanity. Eventually, they discover a path to AGI and engage in intense internal debates: Should they abandon their original mission and
54.
▲
by
WiSaGaN
2y ago
So this is an in-house benchmarks after their undisclosed partnership with a previous benchmark company. Really hope they do not have their next model to vastly outperform on this benchmark in the coming weeks.
55.
▲
by
WiSaGaN
2y ago
H20 is a Hopper GPU, and they are allowed to be sold in China.
56.
▲
by
WiSaGaN
2y ago
Tencent recently bought 100k-200k H20 to serve R1. [1] I think it's not clear open source will tank nvidia price. And you won't place a lot of bets if the outcome is anywhere from certain. [1]: https://aiproem.substack.c
57.
▲
by
WiSaGaN
2y ago
I think the current US administration significantly undermines the framing of democracy versus authoritarianism. As much as Dario may want to push alternative narratives, most AI practitioners—who have no financial ties to OpenAI or Anthrop
58.
▲
by
WiSaGaN
2y ago
Fixed! Thanks!
59.
▲
by
WiSaGaN
2y ago
I think this is a negotiation tactic that aims to eventually block OpenAI's transition to for-profit.
60.
▲
by
WiSaGaN
2y ago
They won't after OpenAI used transformer and refused to give back.
More ›