13 ms·
Anyone else get the feeling this is malicious compliance on behalf of the insurance companies? "Oh, they're going to force us to publish our prices are they? W
by RhodesianHunter 4y ago
Anyone else get the feeling this is malicious compliance on behalf of the insurance companies?
"Oh, they're going to force us to publish our prices are they? Well we'll publish so much data it'll take a herculean effort to make it readable to anyone that doesn't work in data engineering"
- axus 4y agoThe article mentioned CSV files, it seems more like a reflection of what a huge bureaucracy the US healthcare system is. I liked their suggestion that the government should have created the database as part of the law, done the processing on the raw data, and made it more accessible.
- lumost 4y agomeh, I'd prefer the raw data. We can always create DBs out of the raw data, we can always link data. Handling this after the fact would be impossible. Linking a few trillion records doesn't seem that difficult. It should be doable with a good data warehouse and a reasonable entity linking model. I suspect that we'll find more than a few instances of fraudulent behavior once the data is linked. My father was nearly pushed into ~2 Million dollars worth of brain surgery that was unnecessary. Not only was the procedure unnecessary, the price for it was >5X what a top-3 hospital would have charged. I only became privy to this once I pushed him to come to Mass General Hospital (MGH) for a second opinion. The surgeon we saw at MGH also believed the suggested procedure to be dangerous. I wonder if it's possible to cross-reference mortality/complication rates with prices...
- A4ET8a8uTh0 4y agoIn their defense, if it was anything but CSV files, they would be accused of making it too complicated/locking into proprietary formats and so on. I can't say CSV would be my first choice, but I don't really want to think what the alternative would be.
- Victerius 4y agoMillions of .xlsx files
- easrng 4y agoNDJSON? Sqlite?
- ironick09 4y agocsv is many times over again a simpler format than sqlite and easier to understand by anyone across the world. Never heard of ndjson, can’t see one publishing this data in a format that isn’t nearly as common as something like csv (or regular json which some of the data is published in).
- easrng 4y agoCSV only seems simple. Lots of parsing edge cases. Sqlite isn't readable by hand but it's basically bulletproof. NDJSON is literally just newline seperated JSON, it's just easier to process as a stream without a special parser.
- irrelative 4y agoSure, could be simpler, but there's a spec and multiple implementations: https://www.rfc-editor.org/rfc/rfc4180.html https://www.rfc-editor.org/rfc/rfc4180.html And the spec is ~5 pages. Not sure why NDJSON is considered simpler, as json objects can be arbitrarily nested. Breaking into records is easier, but parsing is harder.
- easrng 4y ago(ND)JSON is simpler because people actually follow the spec, unlike with CSV.
- Dylan16807 4y ago
- orangepurple 4y agoimport a CSV into Postgres with open(filepath) as fd: first_line = fd.readline() cols = [] for col in first_line.strip().split(','): col2 = f'''"{col.strip('"')}" text''' cols.append(col2) cols2 = ','.join(cols) print(f"create table {table_name} ({cols2});") print(f"\copy {table_name} from '{filepath}' csv header;") this variant will ingest whatever trash is in your CSV fields as-is (cast & cleanup later) run the output in a psql instance connected to your db (important note: \copy is a psql client command and it is critical to use \copy instead of COPY in many cases where the server process may not have the permission to read your CSV file. with \copy you can read any file the user that launched psql client has permission to read. to make things more confusing it is indeed possible to stream stdin through psql but you use the regular COPY for that instead of \copy)
- RhodesianHunter 4y agoLoading up 100TB of csv data into a Postgres db is not a realistic proposition.
- nardi 4y agoDo NOT read a CSV file by splitting on commas. Python has a perfect CSV library built right in: https://docs.python.org/3/library/csv.html https://docs.python.org/3/library/csv.html. If you split on commas, your code will fail for quoted fields with commas in them.
- orangepurple 4y agoThis is obviously a quick and dirty hack. It only splits the header by commas. Postgres has its own read in logic you can adjust by passing arguments to \copy or copy. If you have anything more complex you need to handle it by parsing the header more intelligently than splitting on commas and adding appropriate arguments to copy or \copy
- deleted 4y ago[deleted]
- daniel-cussen 4y ago
- secondcoming 4y agoSo I should never visit a doctor in Chile?
- daniel-cussen 4y agoI was talking about American doctors, Chilean doctors do wash their hands.
- jjulius 4y ago... what in the blue hell...?
- vxNsr 4y agoMost likely a poor ai attempt.
- 1MachineElf 4y agoFrom the user's "about" section: Perhaps you'll think my comments are unthinkable. My only response to that is that they were legibly written, not by a machine, but by a writer with a soul.
- jjulius 4y agoYeah, it's unclear to me. A lot of the more personal things that are mentioned throughout the account's posts seem to match up with some of the quickly-googleable details that can be found just via their username. I suppose that it could be baked into the AI, but... /shrug
- daniel-cussen 4y agoI got a little excited, that's all I can say about the flow. But there is an undercurrent of betrayal, there absolutely is, undermining everything, only visible if you get trapped in it, if not it looks like a meaningless whoopsie.
- outside1234 4y agoOh totally. It also probably is the formats they already had -- so they just dumped them into a file -- versus making something more orthogonal and ergonomic.
- sl-dolt 4y agoI'm not sure what format they store their records in, but I have a hunch it's a lot more structured than what we see in the CSV files. The data dumps have to comply with some CMS guidelines set out here: https://github.com/CMSgov/price-transparency-guide https://github.com/CMSgov/price-transparency-guide
- kube-system 4y agoThey use relational databases. Then a zillion ETLs to massage that data into every format they need it in, of which this is one of them.
- carabiner 4y agolol, have you ever worked with data from a non-tech company? This is probably the best they have, even inside the company.
- watwut 4y agoCan confirm. Also, it is not better in tech companies, they just have the same data in higher variety of formats and storage systems.
- nerdponx 4y agoNot only was this probably a best effort, but I would bet that at least a few people were busting their asses trying to pull this data together and clean it up.
- teeray 4y agoJust be glad the lawyers didn’t make the prices exclusively available via the traditional UHaul full of Banker’s Boxes.
- geraldwhen 4y agoThere are a lot of billing codes. It’s not as simple as you hope. A giant csv export is easy enough to process and synthesize for normies.
- spaetzleesser 4y agoThey may have that thought but crunching large amounts of data is not exactly hard these days. Better too much than too little data.
- ecommerceguy 4y agoKeep in mind these files are called "Machine Readable" for a reason. Yes, they dumped everything. Machines and 3rd parties are supposed to sort it out. Best be assured each carrier now has full visibility of each other leading to "transparency". It might be a little exciting to be an underwriter right now :D
- ilaksh 4y agoOf course it is. The problem is that the health care costs situation results in many deaths and very severe economic consequences for much of the country. Until lying, cheating, and scheming, and screwing over the public have consequences like prison time, you can expect executives to do everything possible to avoid complying with the spirit of laws like this. There probably was an effort to create a more useful and sanely worded law that would provide a uniform format for rules that could reduce dataset sizes by a factor of 100, but was killed by the healthcare industry because it would require some implementation costs on their end and make the data files actually useful.
- smsm42 4y agoMore like "how we do it in the cheapest way with minimum possible effort". It doesn't earn them any money, so they don't want to spend money on it.
- CogitoCogito 4y agoThey should be forced to provide the costs of procedures up front. As in the whole procedures not just provide prices for a million sub-procedures that the _might_ bill you for. This is how it works in basically every industry. Healthcare tries to argue that it's impossible for them to do this, but is just a load of bullshit.
- danjc 4y agoCame here to say the same thing. I would expect their internal systems to calculate certain pricing components on the fly so it feels like they’ve deliberately built all possible permutations and data dumped that. Basically a document dump - https://en.m.wikipedia.org/wiki/Document_dump https://en.m.wikipedia.org/wiki/Document_dump