7 ms·
I'd like Caltrain to publish raw train data
- guard-of-terra 12y ago"But that’s just the planned schedule" Why won't it match the real schedule?
- bowenli 12y agoCaltrain is often behind schedule. Trains have break down or hit cars. It's a huge pain for daily Caltrain commuters. See: https://twitter.com/Caltrainstatus https://twitter.com/Caltrainstatus
- ak217 12y agoAmong the things Caltrain has to contend with (aside from old equipment prone to breaking) are several dozen at grade crossings and freight train traffic on the same line (!)
- bdamm 12y agoFortunately the freight traffic is mostly after commute hours. Imagine what will happen with those so-called "high-speed" trains coming through!
- guard-of-terra 12y agoAren't you supposed to have separate high-speed track for high-speed trains?
- jarek 12y agoWhile I feel for the grade crossings, freight traffic does not necessarily prevent good service. To give one example, North London Line of the London Overground has freight trains in between every-15-minutes service outside peak hours.
- digitalchaos 12y agoThis is just the exceptional delays. People have given up reporting the "normal" delays we are now seeing every day due to overcrowding of the trains. The overcrowding slows down the onboarding/offboarding of trains by a lot. Another crowdsourced caltrain twitter account is https://twitter.com/caltrain https://twitter.com/caltrain You can see some of the more granular delays there. All these crowdsourced status accounts should be proof that caltrain SHOULD publish the raw data for us to use.
- markcerqueira 12y agoEven aside from major issues like hitting cars or people, delays are not "exceptional" when it comes to Caltrain -- it is the norm. An exception is Caltrain arriving on time.
- britta 12y agoEvery day on Caltrain includes unexpected delays, and it's interesting to look at how those delays affect the system. Check out the linked visualization of Boston-area trains: http://mbtaviz.github.io/ http://mbtaviz.github.io/ - a single day had two major train problems. Caltrain also runs extra trains for special events such as baseball games. Those are planned, but not on the regular schedule.
- guard-of-terra 12y agoMaybe they should be fixing that instead of providing raw data API? Or at least before. Totally doable I should say.
- toomuchtodo 12y agoAwesome! Get a gig with Carltrian and try to eliminate those unexpected delays.
- GFK_of_xmaspast 12y agoMany of Caltrain's problems are systematic, such as lots of at-grade crossings and mega-rich fucks in Atherton.
- guard-of-terra 12y agoJust curious how "mega-rich fucks in Atherton" prevent railway from functioning?
- ZanyProgrammer 12y agoNIMBYs preventing HSR, which in turn jeopardizes the funding course for Caltrain electrification?
- kordless 12y agoMy advice is to never ask someone who is making blaming statements a question regarding the blaming statement.
- superuser2 12y agoBecause reality never matches plans like that. For starters, there is track maintenance, variable loading/unloading time, momentary delays or speed reductions in order to maintain separation when there are a lot of trains running at the same time, and the simple fact that drivers are not necessarily hitting exactly the same acceleration and deceleration curves every single time. By the end of the day, you're going to be more than a few minutes off from how you started. Nonstop long-distance rail service can exactly match its schedule because there are relatively few places for entropy to creep in - you just hold a constant speed across miles and miles of track that you pretty much have to yourself. A commuter rail system is much more complex and there is much more room for entropy.
- snogglethorpe 12y agoIt's certainly not trivial to closely stick to a schedule (e.g. 95% of trains within a minute or two of scheduled times), but it's obviously possible for an urban railway to do so, because many do, often with far more aggressive schedules and higher volumes than Caltrain has. Of course, many American systems are operating at a pretty severe disadvantage, being hamstrung by poor equipment and infrastructure, a lack of funding, understaffing, a hostile political environment, and even pervasive cultural attitudes that dismiss railroads as being something worth investing in. I suppose given all that, it's a wonder they do as well as they do... But still, I think it's important to never forget: it's absolutely possible to do much better.
- superuser2 12y agoThe point is that it doesn't matter. Why should a transportation authority move heaven and earth to satisfy some moralistic concern about keeping exactly to schedules? You don't use CTA (Chicago transit) based on schedules; you go to your stop, read the board with accurate realtime predictions, and wait for a train to show up. It doesn't matter whether the schedule has anything to do with reality, just that you don't have to wait too long. If you'd prefer not to spend too much time on the platform, then you can check your phone, which is accessing the realtime prediction feed anyway, to figure out when you should head to the station. No part of normal usage of CTA depends in any way on the preset schedule, so spending money to keep to the schedule would be waste. Not only is it a Hard Problem, it's one that can be worked around very easily by providing realtime train location and prediction data.
- rakoo 12y agoAuthor, you should integrate your scraper into http://raildar.fr http://raildar.fr, they've already started to scratch that kind of itch for a similar problem.
- tatsiana 12y agoWe've been working on the solution to this issue since our office is overlooking the tracks. You can read more here: http://svds.com/post/listening-caltrain http://svds.com/post/listening-caltrain and here: http://svds.com/post/railroad-modeling-hadoop-scale-hadoop-summit-2014-san-jose http://svds.com/post/railroad-modeling-hadoop-scale-hadoop-s...
- mcpherrinm 12y agoHeh, the author of this post's twitter says he works at Matasano, which appears to be downstairs from you. I guess working next to the tracks has that affect on people!
- bduerst 12y agoThat's pretty cool. I love the idea of scraping real-world information. I did something similar using a GoPro and computer vision when I lived next to one of the 101 off-ramps in SF. I got it to work for most daylight hours (headlights screwed it up hardcore) before our landlord raised our rent by $1000/mo and we moved. I figured it could have been a way to calculate ad impressions for billboards, but I also figured Clear Channel probably already knows those numbers.
- winslow 12y agoI wouldn't be so sure that Clear Channel knows those numbers. You'd be surprised how little data big companies know. Do you have a github or blog post on your experiment?
- deleted 12y ago[deleted]
- ZanyProgrammer 12y agoHeh, I saw this on Twitter and responded to the author earlier-I'm working on a data mining project now with public transit times, comparing arrivals vs scheduled times. Since I live in the Bay Area, it made sense to use local data. However, 511.org, the repository (it seems) for all Bay Area transit APIs, doesn't publish any specific vehicle/route number, or what the actual scheduled time is for an arrival at a stop (though MUNI used to have a nextbus API that was really nicely detailed-I can't find any public hosting of it anymore though). My solution, since I didn't want to do any screen scraping or make trying to identify individual busses/trains a project in and of itself, was to use Portland's TriMet API. That API acutally return specific route numbers, and estimated and scheduled times for each stop (interpolated in the case of non time points). I'm originally from the Portland Area, so I'm pretty familiar with the geography and roads. From what I remember in the 511.org Google developer group, people have raised this exact issue, i.e. Caltrain train numbers. The guy responding from the MTA said they'd try and integrate it in the future, but these posts were like back in 2012 (IIRC).
- simoncion 12y agoIf you're still interested in doing Muni data mining, you'll probably be interested in this: http://www.nextbus.com/xmlFeedDocs/NextBusXMLFeed.pdf http://www.nextbus.com/xmlFeedDocs/NextBusXMLFeed.pdf NextBus is the source of bus position and predicted arrival times for MUNI, and appears to be the same for many other transit agencies. I can verify that (as of three minutes ago) it's still returning reasonable data. However, if you're looking for the SFMTA schedule [0], I don't think you can get it through the NextBus API. I do know that you can get it through a GTFS "feed" found here: http://sfmta.com/about-sfmta/reports/gtfs-transit-data http://sfmta.com/about-sfmta/reports/gtfs-transit-data Also, you might be interested in this, if you haven't seen it already: http://bdon.org/transit/ http://bdon.org/transit/ (SF MUNI transit delays. [This isn't my work.]) [0] Why would you want MUNI's schedule? It's not like any of the drivers care about it! ;)
- ZanyProgrammer 12y agoIt'd be neat if they published positional data. I know the old Nextbus public API for MUNI did that, and it was cool making maps of real time positions of vehicles. I'm sure the excuse now is security BS.
- tzm 12y agoI'd like Caltrain to accept mobile payments.
- enos_feedler 12y agoUse a clipper card? What is the pain?
- tzm 12y agoYes, I have a few clipper cards tied to travel bank accounts. Unfortunately, Clipper cannot be integrated into third-party vendors / apps and is prone to 24 hour account locks if transactions are declined. Adding money ad-hoc is troublesome as well.. use a POTS terminal, go to an approved retailer (Walgreens, etc), online ('available within 3-5 days'). Their commerce system is not mobile friendly and is a pain for mobile users. It could be much more efficient.
- deepsun 12y agoSide note: instead of buying Burp Suite, check out just pure free Chrome or Firefox browsers to watch your HTTP traffic -- they both have pretty good Developer Tools, even IE does. They will show you the returned HTML formatted, and let you change it.
- lstamour 12y agomitmproxy, Charles Web Proxy and Fiddler have also worked for me. I've never understood why someone would pay so much more for Burp if they're not going to use much of it. And for half of the rest, there's plenty of other tools or scripting languages you could use and save yourself a pile of money. I'd love to be convinced otherwise, after all I openly admit I haven't used Burp yet...
- hansnielsen 12y agoI use Burp (and much of its featureset) every day at work; that's why I used it here. The free edition does basically everything you could want except for saving / restoring states (request history, requests you modified, etc). I've also used Fiddler to great effect in the past, but the fact that Burp is written in Java makes it really convenient to use when you deal with multiple OSes.
- bfung 12y agoI also had this idea, but I never executed it as I haven't thought of a way to solve the real vs. estimated times perfectly. Probably can get close w/some data mining, but not sure if it's worth the effort. RE: scraping - instead of putting logic in your scraper, just download the entire section you need, store it in file format. Then parse and shove into database whenever you feel like it. You could rerun the parsing since you'll have all the historically scraped website data on disk.