2 ms·
I think this is a very interesting point. How important is it to be something of a generalist, who could answer questions like this at least in broad terms, as
by defined 9y ago
I think this is a very interesting point.
How important is it to be something of a generalist, who could answer questions like this at least in broad terms, as opposed to being purely a specialist who has little knowledge outside of their domain of expertise?
Would it be significantly helpful to understand more about the things that happen under the hood, as it were, of technologies that we use?
Would a developer who could at least kind of answer this question show an aptitude for broader thought and a deeper interest in technology, and be a potentially more valuable hire as a consequence?
EDIT: To attempt to answer your question, any collection of structured data, whether objects in a memory-based data structure, delimited text fields, or a DBMS file, could be considered a database.
In many data structures, such as hash tables or primitive key/value stores, there is only one key. If you want to find data based on a field that is not the key, you either have to search sequentially through all the records, finding matches on that field, or create an index on that field.
If the number of records is small and the storage medium is fast, a sequential search may be adequate. If not, an index is needed.
Creating an index generally involves scanning all the records in the database and extracting the field required for searching, together with the location of the record within the database. The location would preferably be a direct record number to avoid unnecessary indirection, but it could also be the primary key of the database.
The list of key values and locations is put into a suitable lookup data structure. This could be something as simple as a sorted list in memory, a hash table, or a disk-based structure like a B-tree, B+-tree, or one of many others.
In the most simplistic case, looking up a record using the index means searchng the index for the matching record locator, then using that to retrieve the actual record in a separate step.
Obviously this is a bit more complex for non-unique keys, but that's the general idea.
Finally, the choice of index structure has tradeoffs, because once the index is added, it must be maintained when records are added, deleted, or modified in a way that affects the index. If the db has 100 million records, having to add a new one to a simple sorted index and re-sort it could be a performance disaster.