7 ms·
Disturbing. I did not know SQLite allows writing data that does not match the column type. Yuck. Now I need to review anything I built and fix it. I understan
by andrewstuart 6mo ago
Disturbing.
I did not know SQLite allows writing data that does not match the column type. Yuck. Now I need to review anything I built and fix it.
I understand why they wouldn’t, but STRICT should be the default.
- lateforwork 6mo agoSTRICT has severe limitations, for example it does not have date data type. Why is it a problem that it allows data that does not match the column type? SQLite is intended for embedded databases, where only your application reads and writes from the tables. In this scenario, as long as you write data that matches the column's data type, data in the table does match the column type.
- subhobroto 6mo ago> Why is it a problem that it allows data that does not match the column type? SQLite is intended for embedded databases I'm afraid people forget that SQLite is (or was?) designed to be a superior `open()` replacement. It's great that modern SQLite has all these nice features, but if Dr. Hipp was reading this thread, I would assume he would be having very mixed feelings about the ways people mention using SQLite here.
- SQLite 6mo agoNo, I think that people can use SQLite anyway they want. I'm glad people find it useful. I do remain perplexed, though, about how people continue to think that rigid typing helps reliability in a scripting language (like SQL or JSON) where all values are subclasses of a single superclass. I have never seen that in my own practice. I don't know of any objective research that supports the idea that rigid typing is helpful in that context. Maybe I missed something...
- lateforwork 6mo ago> where all values are subclasses of a single superclass I don't understand this. By values do you mean a row (in database terms)? I don't understand what that has to do with rigid typing. Lack of rigid typing has two issues, in my opinion: First, when two or more applications have to read data from a single database, lack of an agreed-upon-and-enforced schema is a limitation. Second, when you use generic tools to process data, the tools have no idea what type of data to expect in a column, if they can't rely on the table schema.
- subhobroto 6mo agoFirst off, I am so glad the famous "HN conjure" actually worked! My "if Dr. Hipp was reading this thread" was tongue in cheek because on HN it was extremely likely that's precisely what would happen. Thank you for chiming in, Dr. Hipp - this is why I love HN! So, in case you missed it, you're responding to Dr. Hipp himself :) > I don't understand what that has to do with rigid typing. Now I would like to learn a bit from Dr. Hipp himself, so here's my take on it: Scripting languages (like my fav, Python) have duck or dynamic typing (a variation of what I believe Dr. Hipp, you specifically call manifest typing). Dr. Hipp's take is that the datatype of a value is associated with the value itself, not with the container that holds it (the "column"). (I must say I chose the word "container" here to jive with Dr. Hipp's manifest. Curious whether he chose that word for typing for the same reason! ) - In Python, everything is fundamentally a `PyObject`. - In SQLite, every piece of data is (or was?) stored internally as a `sqlite3_value` struct. As a result, a stack that uses Python and SQLite is extremely dynamic and if implemented correctly, is agnostic of a strict type - it doesn't actually care. The only time it blows up is if the consumer has a bug and fails to account for it. Hence, because this possibility exists, and that no objective research has proven strict typing improves reliability in scripting environments, it's entirely possible our love for strict types is just mental gymnastics that could also have been addressed, equally well, without strict typing. I can reattempt the "HN conjure" on Wes McKinney and see if this was a similar reason he had to compromise on dynamic typing (NumPy enforces static typing) to Pandas 1.x df because, as both of them are likely to say, real datasets of significant size rarely have all "valid" data. This allows Pandas to handle invalid and missing fields precisely because of this design (even if it affects performance) A good dynamic design should work with both ("valid" and "invalid") present. For example: layer additional "views" on top of the "real life" database that enforce your business rules while you still get to keep all the real world, messy data. OTOH, if you dont like that design but must absolutely need strict types, use Rust/C++/PostgreSQL/Arrow, etc. They are built from the ground up on strict types. With this in mind, if you still want to delve into the "Lack of rigid typing has two issues" portion, I am very happy to engage (and hope Dr. Hipp addresses it so I learn and improve!) The real world is noisy, has surprises in store for us and as much as engineers like us would like to say we understand it, we don't! So instead of being so cocksure about things, we should instead be humble, acknowledge our ignorance and build resilient, well engineered software. Again, Dr. Hipp, Thank you for chiming in and I would be much obliged to learn more from you.
- andrewstuart 6mo ago>> but if Dr. Hipp was reading this thread He is.
- subhobroto 6mo agoIf you reached out and notified him, Thank you. I hope he has time to revisit - had a few more followups. Cheers!
- andrewstuart 6mo agoNo I did not I think he’s been a regular community member a long time he probably just saw it on front page.
- andrewstuart 6mo ago>> Why is it a problem that it allows data that does not match the column type? “Developers should program it right” is less effective than a system that ensures it must be done right. Read the comments in this thread for examples of subtle bugs described by developers.
- lateforwork 6mo ago> “Developers should program it right” is less effective than a system that ensures it must be done right. You're right, of course. But this must be balanced with the fact that applications evolve, and often need to change the type of data they store. How would you manage that if this is an iOS app? If SQLite didn't allow you to store a different type of value than the column type, you would have to create a new table and migrate data to a new table. Or create a new column and abandon the old column. Your app updates will appear to not be smooth to users. So it is a tradeoff. The choice SQLite made is pragmatic, even if it makes some of us that are used to the guarantees offered by traditional RDBMSs queasy.
- subhobroto 6mo ago> I understand why they wouldn’t, but STRICT should be the default. No wait, what do you mean? As I mentioned at https://news.ycombinator.com/item?id=47619982 https://news.ycombinator.com/item?id=47619982 - your application layer should be validating the data on its way in and out. I mention the two reasons I use for DB fall back
- SQLite 6mo agoChecking the datatype is not the same as validating. There is lots of data out there that is invalid, and yet still has the correct type. In fact, that is the common case. I dare say you will be hard pressed to find a dataset of significant size that doesn't have at least one invalid entry somewhere. Increasingly strict type rules will not fix that.
- subhobroto 6mo agoDr. Hipp, > I dare say you will be hard pressed to find a dataset of significant size that doesn't have at least one invalid entry somewhere I agree. In my experience, and you've forgotten more than I have learned, the mark of a good data engineer is how they account for invalid entries or whether they get `/dev/null`ed. > Checking the datatype is not the same as validating. There is lots of data out there that is invalid, and yet still has the correct type. In fact, that is the common case. Increasingly strict type rules will not fix that. I am having a hard time letting go of this opportunity to learn from you, so in case you have time and read this again - When you say "There is lots of data out there that is invalid, and yet still has the correct type", I read "type" as "shape" or "memory layout" and "invalid" as "semantically wrong". So, is a good example of this a value of `-1` for a person's age? The database sees a perfectly valid integer (the correct shape), but the business logic knows a person cannot be negative one years old (semantically invalid). In that case, to be explicit, `0` or `14` is a "valid type" for an age (usually integer), but completely invalid data if it's sitting in an invoicing application for an adult-only business? Again, thank you for your time and attention, these interactions are very valuable. PS: I'm reminded of a friend complaining that their perfectly valid email would keep getting rejected by a certain bank. It was likely a regex they were using to validate email was incomplete and would refuse their perfectly valid email address
- sgbeal 6mo ago> STRICT should be the default. Somewhat ironically[^1], if STRICT were suddenly made the default, countless applications which work today would fail to work right after that update. SQLite is on billions upon billions of devices, frequently with many installations on any given device, and even a 0.001% regression rate adds up to many, many clients. One of the reasons people trust SQLite is because, as a project policy, it does not pull the rug out from under them by changing long-held defaults. [^1]: it's ironic because proponents of Strict tables frequently say that it improves app stability and robustness, whereas activation of Strict tables in apps not designed to handle it will, in fact, make them more fragile. Spoiler alert: most SQLite apps are not designed for Strict tables.