4 ms·
It's not quite that simple. The CMS detector (one of the two main detectors/experiments operating at the LHC) spits out about 10TB of raw data per second. This
by Analog24 8y ago
It's not quite that simple. The CMS detector (one of the two main detectors/experiments operating at the LHC) spits out about 10TB of raw data per second. This is more than any system can handle at the moment. To deal with this overwhelming flood of data there are layers of "triggers" that effectively filter out the collisions that aren't of any interest. The lowest level triggers are actually embedded in FPGAs on the detector itself, then there is a high-level trigger that runs on an on-site server farm that does minimal event processing to see if the even is interesting.
All of this requires us to define what constituents an interesting event. The most basic trigger is called the "minimum bias trigger" because it requires us to make the fewest assumptions about what makes the event interesting. So even at the most general level a brute force approach will still have some bias because it will likely be using the minimum bias data. What makes it "brute force" is that they will no longer be looking for signals from specific new models, they will only be considering the Standard Model and essentially looking for outliers. The difficult part is that there is so much systematic uncertainty given the complexity of these detectors, which makes it hard to obtain statistical significance.
- hinkley 8y agoOh! So we are using the term brute force in the information theory sense of the word, not the scientific connotation?
- DoctorOetker 8y agocould you elaborate a bit on the "10 TB of raw data per second" ? If it's really raw how much room is there for domain specific lossless compression? How many seconds long has or will the experiment (CMS) normally run? How much did or will the CMS component of LHC cost from construction to its full lifespan? What fraction of that imaginary lump sum went to digital storage budget?
- Analog24 8y agoThese experiments typically run 9 months out of the year with (very roughly) 50% up time. That's (again, very roughly ) 25,000,000 seconds. That's 250000 petabytes a year [0]. And by "raw" I'm talking about the ADC signals from the detector hardware. For example, the readout from the 75 megapixel "camera" (i.e. the tracker) at the core of the detector that takes 40 million pictures a second. When doing analysis the collisions are reconstructed into "events" where all the ADC signals are combined to form particle tracks and energy deposits (allowing us to actually study the physics behind the high energy interactions). These reconstruction algorithms are very complex and statistical in nature so there is some bias associated with them, hence the reason the "raw" data is permanently saved along with the reconstructed data. Compression is most definitely used for the data that is actually saved to disk but it is not done on the fly when trigger systems have to run in nanoseconds. The data limitations are not really from a lack of foresight or proper budgeting. It was known decades ago how much data would be produced by the LHC, it just really is an insane amount of data. [0] after applying triggers (i.e. filtering) only a few hundred petabytes of data is stored in a given year from the CMS detector.
- toomuchtodo 8y agoIt’s been ~10 years since I was at Fermilab working on the CMS Tier 1 team; when I was still new on the team, I expressed concern about data loss due to scheduled maintenance on our Nexsan clusters (coming from the private sector) and my boss would chuckle and say “Don’t worry, they’re always making more data”. Also, scheduled maintenance windows during the day! As an ops guy, I was spoiled for that period of time in my career.
- DoctorOetker 8y agoI'm pretty sure I still don't understand. I'm trying to understand your numbers, at 10TB/s and 40 million pictures per second, thats 250kB per picture. since there are 75 megapixel (or I assume 75 megasensors) how does that work for raw data? Is the energy determined from calorimeters, or from curvature in magnetic field? What exactly is the ADC measuring? deposited charge? Since you are acquainted with the project, could you point me to a good pdf describing the detector and low level digitization architecture (pre FPGA) architecture?
- DoctorOetker 8y agoI assume that not all 75 megapixels are considered hit, and only those hits are digitized by ADC? What is the bit depth of the ADCs?
- Analog24 8y agoYou're reading a bit too much into the "raw data" part. An average event that is saved to disk is about 1MB in size but that is with compression. The uncompressed ADC counts obviously would add up to a lot more than that given the number of readouts in the tracker, which is only one of about a dozen sub-detector systems. So saving every event to disk, with compression, would amount to about 40TB/s. Since there are a number of caveats I was only using an order of magnitude to illustrate the scale of the datasets being handled at the LHC. I'm sorry if I sound dismissive here, I don't mean to be it's just impossible to go into all the detail of these machines in an HN comment. They are unbelievably complex (and fascinating!). I encourage you to check out the CMS design doc that will cover most of it in extreme detail (it's almost 400 pages) [0]. There is also a much more general overview of the various detector subsystems that is considerably more concise [1]. To answer one of your questions though: the tracker is used to measure the track of charged particles through the magnetic field and determine their momentum. This isn't enough to tell us their energy though b/c we don't know how massive each particle is. This is why the tracker is surrounded by an electromagnetic calorimeter (which "catches" electrons and photons) that is surrounded by a hadronic calorimeter (which "catches" hadronic patricles like neutrons and protons). The information from these systems is combined to ID each particle and determine how much energy it had, whether or not it decayed, etc. This is generally how we study the physics involved in the collisions but in reality it is 100x more complicated than this and there are number of other detector subsystems involved. [0] http://inspirehep.net/record/796887/files/fermilab-pub-08-713.pdf http://inspirehep.net/record/796887/files/fermilab-pub-08-71... [1] http://cms.web.cern.ch/news/detector-overview http://cms.web.cern.ch/news/detector-overview
- pasbesoin 8y agoThank you. I find your explanation enlightening and interesting. I haven't spoken with him in some years, and I won't mention names, but a good friend worked on calibration of the CMS detectors (scintillation). The last time I did speak with him, that was a very interesting 2 or 3 hours. Although I don't recall us discussing the data processing aspects in any depth; my limited impression of those I mostly gained from random articles and whatnot. P.S. I guess I'd bettter RTFA, now. ;-)