Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
squarecog
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
The what and why of product experimentation at Twitter
(blog.twitter.com)
13 points
by
squarecog
11y ago
|
7 comments
32.
▲
by
squarecog
11y ago
S3 has a number of challenges of its own, for example, file access is quite expensive (performance-wise). The Netflix engineering team has a ton of experience running data processing "in the cloud", and have their share of war sto
33.
▲
by
squarecog
11y ago
[disclaimer: I managed the data platform group, of which the storm/heron team was part, when heron was developed, and I'm still a Twitter employee; however, naturally, this is all my opinion and not an official position of my empl
34.
▲
by
squarecog
11y ago
Good post explaining table stakes for reasoning about whether your site search works (and measuring whether changes make it work better). If anyone is interested in having this but not rolling their own, www.switfype.com (ycombinator 2012)
35.
▲
by
squarecog
11y ago
Well, first, a lot of this stuff is actually written in Scala, which is like Java's younger and more fun cousin. Still runs on the JVM but has many fun language aspects I won't get into because there's already a billion post
36.
▲
by
squarecog
13y ago
As pointed out by another commenter, Cassandra is not a column store. Its storage format is quite inefficient for compression and for large scans of subsets of columns, which is what columnar stores are designed for (Cassandra and HBase opt
37.
▲
by
squarecog
13y ago
If you are linked to 100 people on your LinkedIn account, and 10 of them have the same set of people in the address book they imported into LinkedIn, it's not hard for LI to figure out you might also know people in that overlapped set,
38.
▲
by
squarecog
13y ago
How do you imagine it being involved?
39.
▲
by
squarecog
13y ago
That's an interesting thought, but rather than focusing on write vs read balance, I would claim that OLAP and OLTP are most distinguished by the nature of queries they need to support. OLAP is characterized by fairly low-volume aggrega
40.
▲
by
squarecog
13y ago
It's not the data coming in that makes things hard. It's the fanout of messages into their subscribers' timelines.
41.
▲
by
squarecog
14y ago
How do you keep the message queue from unbounded growth?
42.
▲
by
squarecog
14y ago
Which CrossFit gym in SOMA are you referring to? (Mine is awesome, too. Just wondering if it's the same one.)
43.
▲
by
squarecog
14y ago
Relevant auto-complete out of the box is a killer feature. Feels and works so much better than the Tumblr defaults, at least based on the reference implementations. Nice job @swiftype.
44.
▲
by
squarecog
14y ago
No -- this is how they weeded out prima donnas.
45.
▲
by
squarecog
14y ago
450 Megs? Use pig -x local :)
46.
▲
by
squarecog
14y ago
50 Megs? How in the world are you getting that? I just checked trunk, and it's 3.1. Granted, Cascading comes in at about half a meg for 2.0, but this 50 meg figure is just silly. Our tests take seconds, and can run from eclipse (we use Pig'
47.
▲
by
squarecog
14y ago
And Twitter. We unit test our pig scripts, it's pretty straightforward given MockStorage class we have contributed (you can find it in pig trunk). Granted, we've long been separating load statements from actual logic, which allows us to fai
48.
▲
by
squarecog
14y ago
We don't use Azkaban or Oozie at Twitter, so we are unlikely to do that. But we'll happily take a good pull request. I think the LinkedIn guys are interested.
49.
▲
by
squarecog
14y ago
JSON is a good choice if you have very tight coupling between producers and consumers of your logs. Twitter (disclosure: I manage/was early engineer on the analytics infra team) used Json initially, and it quickly turned into a mess because
50.
▲
by
squarecog
15y ago
You are doing regex matching in the Cascading code, but splitting on a character in the pangool code. The latter is obviously much faster. I don't know that that's the reason for the difference you observe, but it certainly can't hurt to fi
51.
▲
by
squarecog
15y ago
Most of the "forks" are forks from Twitter employees' personal repos that were created before things got consolidated under Twitter's account. A few aren't and are called out as such. Which ones are missing proper attribution?
52.
▲
by
squarecog
15y ago
Disclaimer 1: I work for Twitter. Disclaimer 2: Not a twitter spokesperson, opinions are my own. I don't understand how so many people are missing the point demonstrated by http://www.focusontheuser.org/ . It uses information available fr
53.
▲
by
squarecog
15y ago
Nothing is wrong with byte buffers. Use them where appropriate. The advice is to stop when you find yourself implementing a full-blown memory manager / quasi-malloc in user code on top of byte buffers...
54.
▲
by
squarecog
15y ago
The context isn't clear from these notes. Full context as explained in the talk: used to have stop-the-world GC for 2 minutes every hour. After implementing bytebuffer-based slab allocation, this is only several seconds, and once every thre
55.
▲
by
squarecog
15y ago
Except, supermodels don't eat.
56.
▲
by
squarecog
15y ago
Couple hours later, Hogan is faster than Handlebars: http://jsperf.com/t-bench2/7
57.
▲
by
squarecog
15y ago
I would argue that any time you put "just" and "terabytes" next to each other, you are heading for big problems to go with your big insights :). Schema-less is great.. until you can't find stuff and your data is full of inconsistencies.
58.
▲
by
squarecog
15y ago
The HBase integration with Pig is pretty good (disclaimer: I wrote a bunch of it, and use it on a daily basis). The only thing is that you need to create the table and set up column families yourself. The mongo driver Russel demoes automati
59.
▲
by
squarecog
15y ago
Where did you get that this is being marketed for production? Did you just assume that because the site looks nice? Handy tip: languages that are marketed for use in the real world tend to advertise actual uses in the real world.
60.
▲
by
squarecog
15y ago
CoffeeScript. CS. JavaScript. JS.
More ›