4 ms·
In my text editor (https://github.com/alefore/edge https://github.com/alefore/edge) I use a balanced binary tree containing small chunks (std::vector) of contig
by afc 3y ago
In my text editor (https://github.com/alefore/edge https://github.com/alefore/edge) I use a balanced binary tree containing small chunks (std::vector) of contiguous lines.
That works well enough for me: https://asciinema.org/a/314752 https://asciinema.org/a/314752
This requires loading the entire file into memory, but computers are so fast today that optimizing that away is pointless IMO (as the video shows, a 12MB file with 400k lines can be loaded/edited/saved reasonably).
The tree structure is implemented as a generic container here: https://github.com/alefore/edge/blob/master/src/language/const_tree.h https://github.com/alefore/edge/blob/master/src/language/con.... The tree is actually deeply immutable; each modification operation returns a new tree (where most contents are shared with the original tree). The leafs contain std::vectors that hold between 128 and 256 lines (save for a few cases). The actual specialization that holds a sequence of lines is LineSequence::Lines here: https://github.com/alefore/edge/blob/master/src/language/text/line_sequence.h https://github.com/alefore/edge/blob/master/src/language/tex...
- vidarh 3y agoI just use a (Ruby) array of strings for mine, on the basis that I have never on purpose wanted to actually edit a file large enough that this is a problem. Computers keeps getting faster, but files I want to edit really don't. If I open a huge one it's usually either an accident or because I want to view it, not edit it. It's easy enough to support abnormally large files, but it felt to me like a distraction that adds complexity for the sake of a fringe case I was perfectly happy to just define as out of scope. Incidentally, the server process for my editor holds every buffer that I've ever opened (over several years) and not explicitly killed in memory, and it's consuming too little for me to bother adding any garbage collection for buffers or similar (it's synced to disk, so on reboot I have the full same set of buffers available). We're talking many hundred buffers. RAM is cheap and plentiful.
- cdcarter 3y agoYou've worked primarily in your own home built editor for years? That's pretty cool.
- afc 3y agoI'm not the person you directly replied to, but I've worked on my text editor since 2014 and used it exclusively since ~2015. It's been a lot of fun!
- vidarh 3y agoYeah. 6 years or so now. Putting the buffers in a server that checkpoints regularly was key to making it easy to do early - made it really hard to lose any data even when it crashed regularly.
- atombender 3y agoHow does that perform when the entire 12MB file is a single line? Before you say it's unrealistic; sure, it's not a normal use case for code, but I frequently find myself wanting to open a huge JSON file, for example, just to perform a search/replace, or a minor edit, or copy some fragment, or just look at its contents. Pretty-printing first, or doing the change with a specialized tool like jq, is also something I do, but being able to just load the file into a text editor is often more convenient. In my experience, editors tend to break down on edge cases like these. Editors (e.g. Jetbrains) often turn off syntax highlighting and/or edit capabilities on large files, not to mention put an upper limit on the number of cursors.
- afc 3y agoThank you for your response. You ask a good question. I expect very long lines should work just fine, but I'll give it a try when I have a chance. I represent lines using a virtual class that has a few implementations. IIRC, the line will start as a tree containing 64kb chunks. As edits are applied, these chunks will gradually get replaced with other implementations (of the virtual class) that delegate to the original instances (and the representation is occasionally optimized (which can incur copying) to avoid too much nesting). So simplifying, I think this should just work due to the optimized representation that I use for lines, which doesn't require the characters to be contiguous in memory. I fully agree with your observation about how things often break in the corner cases. I actually also put a configurable cap on the number of cursors (e.g., a reg exp search stops after 100 matches). For syntax highlighting I don't currently have a cap, but this is ~fine: it just ~wastes a background thread (all syntax parsing is offloaded to a background thread, not blocking the main thread). But yeah, I should probably make this thread give up after a configurable time out.
- JoeyDoey 3y agoLooks like Codemirror does similar with long/ huge files https://codemirror.net/examples/million/ https://codemirror.net/examples/million/
- layer8 3y ago12 MB isn’t necessarily that much. Here is a 96 MB SVG file: https://commons.m.wikimedia.org/wiki/File:Koppen-Geiger_Map_Cfc_present.svg https://commons.m.wikimedia.org/wiki/File:Koppen-Geiger_Map_... Here are genome sequence files that are several hundred MB compressed (the decompressed format is text): https://ftp.ncbi.nlm.nih.gov/genomes/refseq/ https://ftp.ncbi.nlm.nih.gov/genomes/refseq/ Finally here is an almost 2 TB XML file: https://wiki.openstreetmap.org/wiki/Planet.osm https://wiki.openstreetmap.org/wiki/Planet.osm These aren’t the most typical use cases, but it’s nice if an editor can at least handle > 32-bit files (larger than 2 or 4 GB).
- SpaghettiCthulu 3y agoIt's all fun and games until you're loading a multi gigabyte CSV file
- Findecanor 3y agoYou would have to load the entire file either way, even if you'd mmap() it, to be able to break the lines to find their extents and start offsets. mmap() isn't entirely safe under Unix/Linux anyway because there is no Mandatory Locking. And it would still be useful only if you operate on the original text encoding.
- afc 3y agoYeah, good point. You want to at the very least be able to show the user on which line his cursor is, and that already requires reading all the contents from the start up to the cursor. You might as well read the entire file and also show the user the total line count.