Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
srush
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
61.
▲
by
srush
3y ago
Nowhere near as neat as candle or ggml, but just released a 4-bit rust llama2 implementation with simd. Runs pretty fast. https://github.com/srush/llama2.rs/
62.
▲
by
srush
3y ago
There is also a more interactive version if you want to challenge yourself. A Python notebook of interactive puzzles to build an adder with transformers. https://github.com/srush/Transformer-Puzzles
63.
▲
by
srush
5y ago
We fine-tuned the model on a dozens of different NLP datasets and tasks in a prompted style. You can read all the prompts in the appendix or get them all here: https://github.com/bigscience-workshop/promptsource . Most
64.
▲
by
srush
5y ago
While in theory it could, the nature of its training favors shorter more factual replies.
65.
▲
by
srush
5y ago
Yes, this is an important challenge. There has been a lot of interest in the NLP community right now, particularly around QA tasks [1] Standard supervised models do it well, but zero-shot models still have trouble. 1. https://arx
66.
▲
by
srush
5y ago
We discuss this a bit in Section D.2 (HOW UNSEEN ARE THE HELD-OUT TASKS?). From our perspective, a) The tasks we test on are very different, particularly tasks like BIG-Bench that we didn't even have access to until several days ago (a
67.
▲
by
srush
5y ago
Yes there are many reproducible measures for benchmarking NLP datasets. We use many of them in the paper. The issue here is that we were not completely sure of the process that OpenAI used in their paper. They report the prompt but not the
68.
▲
by
srush
5y ago
from @craffel: It's possible to run inference on a single Google Cloud TPU v3-8 device or on a server with 4x 32GB v100 GPUs. Hugging Face also has an inference API for any model on the Hub: https://api-inference.huggingface
69.
▲
by
srush
5y ago
It actually does get it "right" if you fix the typo :)
70.
▲
by
srush
5y ago
The results presented in this paper are for "true" zero-shotting in the literal sense that the model has never been explicitly trained on the tasks presented, nor do we cross-validated on the prompt choice.
71.
▲
by
srush
5y ago
Full information about the BigScience Project is here https://bigscience.huggingface.co/
72.
▲
by
srush
5y ago
Yes. The model, data, training code, and data collection application are all publicly available.
73.
▲
by
srush
5y ago
I've been playing around with some of these ideas for Python notebooks. https://github.com/srush/streambook This is a proof of concept that combines Jupytext for markdown readable notebooks with Streamlit for in-o
74.
▲
by
srush
6y ago
Hi all, I'm one of the original authors of OpenNMT and an nlp prof, always nice to see it trending :). I work with Hugging Face ( https://huggingface.co/ ) now, we do similar things for NLP generally as well as supportin
75.
▲
by
srush
6y ago
Yup that's right. Although I would add front-end search and visualization (with d3) to the features. As you mentioned there are a lot of good solutions for review, notification, and reg. Common choices are OpenReview/CMT/HotC
76.
▲
by
srush
6y ago
Yeah :) we hacked in a lot of things last minute for ICLR. MiniConf is a clean 80% functionality / 20% code version of that codebase.
77.
▲
by
srush
6y ago
Thanks! If you are interested more about the event here was a podcast https://www.thetalkingmachines.com/episodes/iclr-accessible-... and blog post https://medium.com/@iclr_conf/gone-virtual-lesson
78.
▲
by
srush
6y ago
Thanks for the comment! Updated the readme. During the conference we ran we did integrate chat and video tools (Rocket.chat, slido, slideslive). This is really just the glue to pull those parts together.
79.
▲
by
srush
6y ago
Hi HN, I'm the developer (@srush_nlp). This was made for ICLR a deep learning conference we ran last month. We couldn't find any tools to do the things we wanted so we went a bit rogue and built it ourselves. There's a bunch
80.
▲
ICLR Public Archive
(iclr.cc)
1 points
by
srush
6y ago
|
0 comments
81.
▲
by
srush
8y ago
(I'm the author) Thanks this is really helpful. I agree with the points being made. The post started out with the hypothesis, "look this basically works as is", but the conclusion of the UX experiments seems to be "the i
82.
▲
by
srush
8y ago
Sure, happy to respond. > Have you worked with professional translators much along the way? Most of us use and keep our own translation memories stretching into the millions of segments. I personally haven't worked too much with tra
83.
▲
by
srush
8y ago
> Do you rearchitect the code upon new research breakthroughs? Great question. So we now have implementations of Transformer in both the PyTorch and TensorFlow version of the library. It did require some rearchitecting, particularly for
84.
▲
by
srush
8y ago
This is probably true in general, but in the last year with the switch from Torch => PyTorch the core code has actually dropped in size. There is still progress being made in improving the frameworks for specifying deep learning models.
85.
▲
by
srush
8y ago
Hi everyone, I'm a creator of OpenNMT and run the NLP group at Harvard ( http://nlp.seas.harvard.edu ). A lot has changed in both translation and deep learning over the last couple years. Happy to answer any questions about t
86.
▲
by
srush
9y ago
We published some research on one approach: http://lstm.seas.harvard.edu/latex/ Here's how to do it with OpenNMT/PyTorch: http://opennmt.net/OpenNMT-py/im2text.html
87.
▲
by
srush
10y ago
If you are interested in looking at the model in more detail, we (@harvardnlp) have uploaded the model features to LSTMVis [1]. We ran their code on amazon reviews and are showing a subset of the learned features. Haven't had a chance
88.
▲
by
srush
10y ago
Yeah this is a nice connection. Note however that there has been much less success in using perturbation in language. The fact that inputs are discrete makes it harder to apply some of the tricks in the adverserial image work.
89.
▲
by
srush
10y ago
You can often get something started with ~10000 examples. It's very problem specific though.
90.
▲
by
srush
10y ago
Non-standard datasets like these are very fun to play with. If you can produce the training data (source => target aligned text files), it is relatively simple to try it out. Some mappings that people have recently published on: code =&g
More ›