Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vineetg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Biotech Platform Benchling Valued at $6.1B in New Funding
(bloomberg.com)
11 points
by
vineetg
5y ago
|
0 comments
2.
▲
Sqlalchemy_batch_inserts: For when inserting lots of rows is slow
(benchling.engineering)
1 points
by
vineetg
7y ago
|
0 comments
3.
▲
by
vineetg
8y ago
Original author here. Thanks for the kind words! We would love open source some of the work we did - there are a few edge cases to still work out with deprecated_column and renamed_to before I’d be comfortable doing that, but definitely agr
4.
▲
by
vineetg
9y ago
Original author here. You're right - conceptually the CRISPR search problem and DNA sequence alignment are related. In both, you're looking for place where two (or more) sequences are very similar. I would say there are two major
5.
▲
by
vineetg
11y ago
I briefly mentioned it in the "New Infrastructure" section, but we're doing the splitting and combining of results on our web servers.
6.
▲
by
vineetg
11y ago
We're getting around 100ms connection times. We're using the Node aws-sdk to get things from S3: s3.getObject({Bucket, Key})
7.
▲
by
vineetg
11y ago
EBS costs are actually fairly small ($9 a month per instance for 90GB). More than 95% of our costs were just paying for the EC2 servers.
8.
▲
by
vineetg
11y ago
We actually do this in our "expensive" check. The reason that we use the 2 arrays is because in 2 array lookups, we can check for matches for all 200 guides. I probably could've made that clearer in the post - the array conta
9.
▲
by
vineetg
11y ago
A lot of these are actually pretty easy to spot with tools like gcov. To determine false positives, we can just look at how many times each "if" statement was hit, and compare those to the final count.
10.
▲
by
vineetg
11y ago
Thanks for bringing up all of these alternatives. We definitely would have preferred to use an existing solution to building our own. Unfortunately, a lot of the existing software is not intended for the search we're trying to do, or d
11.
▲
by
vineetg
11y ago
We messed around with bowtie - it seems like most of these are optimized for the number of alignments being small (i.e. close to 1). Unfortunately, the number of matches for a 20 base guide on the human genome is closer to 1000.
12.
▲
by
vineetg
11y ago
I left PAM sites out of the blog post (it's actually mentioned briefly in the footnotes), as it made the problem slightly more complicated. The final algorithm actually keeps track of the last 20 bases + PAM length, and checks both the
13.
▲
by
vineetg
11y ago
(Author of the blog post here) We use the scoring function published by Hsu et al[1], which most scientists seem to be using. This function takes into account both the number of mismatches and where they occur in the guide. There's a m