Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
WiSaGaN
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
151.
▲
by
WiSaGaN
3y ago
There was debate about whether to “buy from abroad or to make ourselves” in China. The debate ceased to exist in recent years. Right now, one of the headlines in HN is “US wants ASML to stop servicing China-owned chip equipment”. [1] I can
152.
▲
by
WiSaGaN
3y ago
More and more companies that were once devoted to being 'open', or were previously open, are now becoming increasingly closed. I appreciate Stability AI releases these research papers.
153.
▲
by
WiSaGaN
3y ago
Could the reason that 3 states in this case be more efficient than 2 states be that 3 is closer to 2.718... (Euler's number) than 2 is?
154.
▲
by
WiSaGaN
3y ago
Definitely, should a reputable media (at least ten years ago NYT was reputable) report this complexity and nuances about these people or about the situations so that readers can get a complete view?
155.
▲
by
WiSaGaN
3y ago
I find it facinating how liberal mainstream media (MSM), such as The New York Times (NYT), positions itself as an opponent to xenophobia and corporate dominance in America domestically. However, when it comes to international affairs, it se
156.
▲
by
WiSaGaN
3y ago
openrouter provides access to gemini, and it has openai compatible API too. [1] [1]: https://openrouter.ai/docs
157.
▲
by
WiSaGaN
3y ago
Changelog is also updated: [1] Feb. 26, 2024 API endpoints: We renamed 3 API endpoints and added 2 model endpoints. open-mistral-7b (aka mistral-tiny-2312): renamed from mistral-tiny. The endpoint mistral-tiny will be deprecated in three mo
158.
▲
by
WiSaGaN
3y ago
mistral-medium has been dated and tagged as mistral-medium-2312. The endpoint mistral-medium will be deprecated in three months. [1] [1]: https://docs.mistral.ai/platform/changelog/
159.
▲
by
WiSaGaN
3y ago
I have recently been amazed by the doublespeak used by these media outlets to describe China's economy. China's real GDP growth for 2023 was 5.2%, yet it was labeled as 'sluggish', whereas the US growth rate was 2.5%, bu
160.
▲
by
WiSaGaN
3y ago
Missed that! Thanks for pointing out!
161.
▲
by
WiSaGaN
3y ago
That seems to be per card instead of chip. I would expect it has multiple chips on a single card.
162.
▲
by
WiSaGaN
3y ago
How much do 568 chips cost? What’s the cost ratio of it comparing to setup with roughly the same throughput using A100?
163.
▲
by
WiSaGaN
3y ago
Great article. It was a joy to read. I have one question though: Why do we integrate the control vector across all layers of a neural network, rather than limiting its application to just the final layer or a subset of layers? Given that ea
164.
▲
by
WiSaGaN
3y ago
This is not an LLM, so it doesn't make much sense to compare it with one. A more appropriate comparison would be something along the lines of GPT-4 OS-Copilot vs. GPT-4 CoT or vs. GPT-4 zero-shot.
165.
▲
by
WiSaGaN
3y ago
This is mostly because of seasonality. Chinese new year in 2023 was in January. In 2024 it is in February. Price always rises the biggest in Chinese new year because people tend to consume not produce during the period. Thus Jan YoY will be
166.
▲
by
WiSaGaN
3y ago
So you created this account just to make this comment.
167.
▲
by
WiSaGaN
3y ago
My bad!
168.
▲
by
WiSaGaN
3y ago
It seems the "gpt-3.5-turbo-0125" mentioned in the blog is not available yet through API as of 01-26 01:18 UTC? Using it resulted "The model `gpt-3.5-turbo-0125` does not exist or you do not have access to it.". It is no
169.
▲
by
WiSaGaN
3y ago
There is an issue for this: [1]. I think it's more of priority issue. [1] https://github.com/ollama/ollama/issues/305
170.
▲
by
WiSaGaN
3y ago
Thank you for your answer.
171.
▲
by
WiSaGaN
3y ago
I've come to understand that parents who have lost a child often face significant challenges in maintaining their relationship. Could you offer advice for such parents, both for before and after experiencing this tragic event, on how t
172.
▲
by
WiSaGaN
3y ago
The character '鉞' is not commonly used in modern Chinese. I would seriously doubt that anyone teaching Chinese would use this word to illustrate the letter 'Y' in English or Pinyin, especially to toddlers. There are plen
173.
▲
by
WiSaGaN
3y ago
I’m wondering what cost function it’s using. Does it use a simulator?
174.
▲
by
WiSaGaN
3y ago
In my experience, it's great at its size, but obviously worse than mistral:7b-instruct-v0.2. Currently mixtral:8x7b-instruct-v0.1 is the lowest inference cost model at similar performance level of GPT-3.5.
175.
▲
by
WiSaGaN
3y ago
Although Rust is an amazing language with rapid adoption, it still has a relatively smaller user base in low-level programming, where llama.cpp operates. As a result, the pool of talent that can contribute to such projects would be more lim
176.
▲
by
WiSaGaN
3y ago
Any existing stream api for llm input?
177.
▲
by
WiSaGaN
3y ago
The best analogy I can think of is that if you want your agent to accomplish something without persisting a long chat history, but instead use the agent to reorganize and change the prompt, you would choose to use the method Leonard uses in
178.
▲
by
WiSaGaN
3y ago
I am not sure about that. You are already using API, should be trivial to use a tokenizer to get the number. Also prompt is just minor part of the cost. You have much less control in the completion part, which is the majority of the cost.
179.
▲
by
WiSaGaN
3y ago
China is still in its infancy regarding propaganda. Russia performs better than China, but it still lags far behind the state of the art in both depth and prevalence.
180.
▲
by
WiSaGaN
3y ago
I am wondering why it would price them in characters but not tokens? Are they processing characters directly as tokens without tokenizer?
More ›