22 ms·
We actually put a bunch of effort into consistency/repeatability checks. Every interview is recorded (video), and we re-watch and re-grade a percentage of them
by ammon 10y ago
We actually put a bunch of effort into consistency/repeatability checks. Every interview is recorded (video), and we re-watch and re-grade a percentage of them to measure the consistency. A long-term experiment we're running is comparing qualitative scores (code quality, good process, how good did the interviewer feel the candidate was) with quantitative features (which tests passed, how long did it take, what design--picked from a decision tree--did the candidate take). We calibrate the qualitative scores with the recorded interviews. So far, quantitative scoring is winning (when judged against predicting interview results at companies). We're waiting, however, until we can see which better predicts job success.
- sillysaurus3 10y agoNotably absent from that list is "Are we verifying we're doing a good job as interviewers?" It doesn't matter how good the interviewer feels the candidate is, or whether a design was picked from a decision tree. All that matters is whether the candidate can do the work at actual companies. I think people here are reacting to irrelevancies during the interview process -- questions which cannot possibly be reflective of a candidate's real-world competency. (When was the last time you shifted a gigabyte of memory? And even if you did, that's not what companies are going to employ people to do. So why ask the question? Are you sure it isn't trivia?)
- ammon 10y agoInterviews absolutely should be grounded in trying to predict how a candidate would do on a job. That's the whole ballgame. The question is how to best do that. First, you need to run a repeatable process (my previews comment). Second, you need to look at the right skills. The approach we take is to track a lot, and figure out what works the best over time. What we've found to be most predictive (so far) is a base level of coding competency, plus max skill (how good an engineer is at what they are best at). So (beyond the coding portion) we don't actually care very much about what a candidate is bad at. We care about how good they are at what they are good at. To give as many candidates as possible the opportunity to show strength, we cover a number of areas. This includes back-end web development, distributed systems, debugging a large codebase, algorithms, and -- yes -- low-level systems (currency, memory, bit and bytes). We do not expect any one engineer to be strong in all of these areas (I'm weak on some of them). But they all are perfectly valid areas to show strength (and we work with companies that value each of the areas). We've recently moved to a new interview processes organized around this idea of max skill. It's working great in terms of company matching and predictive ability. However, it seems we may have underestimated the cost to candidates of being asked about areas where they are weak. There's more negative feedback here than we've seen in previous HN discussions, and I think that the interview change may be behind that. I'm taking that to heart. I think we can probably articulate it better (that we measure in a bunch of areas and look for max strength). We're also running an experiment now where we ask engineers are the start of the interview which sections they think they'll do best on. I'm excited about this. If engineers can self-identify their strongest areas, we'll be able to make the process shorter and much more pleasant! So, the bit shift question: that come up down one branch of a system design question that we used for a while (we've since moved to a more targeted version that is more repeatable). The (sub)issue involved adding a binary flag to a large data blob (this came up as part of a solution to a real-world caching problem). Adding a single bit flag to the front of a 1GB blob has a problem. To really add just one bit, you'd have to bitshift the entire 1GB. This is clearly not worth it to save 7 bits of storage (ignoring that that would not be saved in any case). You can just use a byte (or word), or add the flag at the end. When candidates suggested adding a bit flag at the front, we would follow up asking them how they'd do it (to unearth if they were using 'bit' as a shorthand for a reasonable solution, or if they really are a little weak in binary data manipulation). This was one small part of our interview. By itself it in no way determined the outcome of the interview, or even of the low-level systems section. Plenty of great engineers might get it wrong. But I don't think it was unfair.
- sillysaurus3 10y agoSo, the bit shift question: that come up down one branch of a system design question that we used for a while (we've since moved to a more targeted version that is more repeatable). The (sub)issue involved adding a binary flag to a large data blob (this came up as part of a solution to a real-world caching problem). Adding a single bit flag to the front of a 1GB blob has a problem. To really add just one bit, you'd have to bitshift the entire 1GB. This is clearly not worth it to save 7 bits of storage (ignoring that that would not be saved in any case). You can just use a byte (or word), or add the flag at the end. When candidates suggested adding a bit flag at the front, we would follow up asking them how they'd do it (to unearth if they were using 'bit' as a shorthand for a reasonable solution, or if they really are a little weak in binary data manipulation). This was one small part of our interview. By itself it in no way determined the outcome of the interview, or even of the low-level systems section. Plenty of great engineers might get it wrong. But I don't think it was unfair. Of course it's unfair. The candidate isn't actually programming a solution when they're talking to you. They're on a tight time crunch, under a microscope, in front of an interviewer. The answers to your questions will literally make or break their future with you. Did you specify to them in your original question that the entries in the cache are 1GB large? If you assigned them the task of implementing a solution to your caching question, they would immediately notice using a 1-bit flag is a poor design decision. The point is: Plenty of great engineers might get it wrong. That says quite a lot more about Triplebyte's question than the engineers. A wrong answer doesn't mean they're weak in bit manipulation or that they decide to implement poor solutions. It says they're suffering from interview jitters. They're weak in the artificial environment you've constructed for the purposes of the interview, which may or may not correlate with their actual ability. This may sound like useless theorizing, but unfortunately a massive number of excellent engineers are awful in an interview setting. But if you give them a problem to actually solve, they pass with flying colors. Triplebyte does give problems to candidates to solve, but it sounds like you also care about whether they can pass your interview (by demonstrating sufficient max skill when prompted) instead of whether they can implement solutions to the problems you assign to them. This rules out candidates who would otherwise do very well, which is the type of candidate you're trying to find. I know that you're saying the verbal section of the interview isn't the whole process, but are you sure it's an effective one? It might be positively misleading. A candidate who is very strong in the area you're looking for is also likely to be someone who will get your questions completely wrong, because they're not programming. They're talking. So it sounds like you're selecting for people who can talk well: those who can show strength during your interview when prompted verbally. Is that the right metric to find talented candidates? If you were to put together a pipeline where e.g. you give candidates an XCode codebase and say "There are bugs in this codebase, and <specific missing features>. Implement as many fixes or improvements as you wish or have time for, then send us the code," you would have a mechanism which selects for candidates who are ~100% competent, since that's exactly the type of work they'll be doing on a day-to-day basis. Some candidates wouldn't want to do that, so perhaps there should be an alternative for them. But it'd be vastly more effective than quiz-style questions during a timeboxed interview. It's possible to come up with endless reasons why it might be a bad idea to set up a pipeline like that. But all the companies that have set it up have been shocked how well it works when they rely solely on that test. Instead of an opportunity to show strength during an interview, the candidate is able to directly answer the question "Can they do the work?" EDIT: From one of the other comments (https://news.ycombinator.com/item?id=13834231 https://news.ycombinator.com/item?id=13834231): > I applied through their project track. It was described as a low-pressure way to write your code ahead of time and talk about it in the interview. The interview was, instead, about making changes to my project while Ammon watched. (Also, there was a request to derive a formal proof while Ammon watched. I didn't get it.) After which I got a rejection saying that my project was great but my interview performance was so poor that they wouldn't move forward. It sounds like Triplebyte almost has the pipeline described above, but it won't work if you watch the candidate or ask them to do more work. The project alone has to be set up to be a sufficient demonstration of skill.
- mbesto 10y ago> better predicts job success How do you rate job success?
- pyb 10y agoWasn't there a study by Google a while back, where they found, as a trend, that their most successful people had only marginally passed their job interview ?
- astrange 10y agoThat sounds like the Pareto-optimal solution to a job interview. It's like how you're not doing grad school properly if your grades are better than C. Also, candidates who do too well could be too good for the job, since Google supposedly likes having incredibly overqualified people maintain do-nothing internal apps.
- jameside 10y agoThe finding was that people who had received a "no hire" recommendation yet still received an offer tended to do well. The reason being: to compensate for the poor feedback, someone else on the hiring loop believed in the candidate so much and saw something exceptional and were willing to go to bat for them.
- pfarnsworth 10y agoIt sounds like your ability as an interviewer is pretty poor. There are several examples of the same smug behavior. What sort of training have you gone through to ensure that you're actually an appropriate and qualified person to be interviewing?
- dkarapetyan 10y agoI interviewed with Ammon. He did not come across as smug. I went through the process just as a curiosity because I was fed up with the typical interview process. I haven't followed through anything else related to Triplebyte in terms of an actual job in case people think I have a favorable view because I got a job through them.
- Harj 10y agoI'm his co-founder and this comment is unnecessarily personal. Ammon has done over 900 technical interviews (https://www.reddit.com/r/cscareerquestions/comments/5y95x6/i_am_ammon_bartram_and_i_have_done_900/ https://www.reddit.com/r/cscareerquestions/comments/5y95x6/i...) and there's one negative reference to him specifically on this thread.
- pfarnsworth 10y agoNah, you can't use forum threads as proof. You need to do the analysis scientifically, which based on the feedback, it sounds like you need to do. You don't bother quantifying the ability of you and your interviewers, you just assume you're great.