4 ms·
Tangentially relevant: Gernots list of benchmarking crimes. https://gernot-heiser.org/benchmarking-crimes.html https://gernot-heiser.org/benchmarking-crimes.ht
by wucke13 2y ago
Tangentially relevant: Gernots list of benchmarking crimes.
https://gernot-heiser.org/benchmarking-crimes.html https://gernot-heiser.org/benchmarking-crimes.html
- bjackman 2y agoHa, I thought this would be a really useful resource but I think the people the author is complaining about do much better than most benchmarking I see in the industry Almost all the benchmarking results I see is just a percentage difference between two algebraic means, no statistical analysis whatsoever. Very common interaction: QA folks say "your change degraded some of our metrics and improved some others". I know they are full of shit because it's impossible that my change improved any perf metrics. I ask for statistical details, they don't have any, this meeting was a waste of time, it will be next time too. The fact that I get these reactions suggests that everyone else just lets each other get away with it.
- porridgeraisin 2y agoYep. The most recent example that's stuck in my head is actually much worse: they didn't even take the mean! One sample! https://github.com/denoland/pm-benchmark https://github.com/denoland/pm-benchmark Check the run bench shell script (there's not much else in the repo anyways)
- genewitch 2y agoHey, that's perfectly valid for arguing with your friend about which one to deploy on our server, all things equal. I do this sort of thing to see what tools are faster all the time. ripgrep, ag(silver searcher), grep, MongoDB was one we were arguing about for a while recently.
- burntsushi 2y agoWhich one won? :) (I'm the author of ripgrep.)
- genewitch 2y agoripgrep, except against the full 160GB dataset, mongoDB was faster on my ryzen. I have a lot of subtitles. I'm partially hard of hearing and partially i can't stand the way everything is mastered, so i use volume normalization (sometimes called "night mode", vizio calls it this) and subtitles to make up for the fact that the audio tracks in most things is bad. Well a side effect of subtitles is now i have context for every video that i can search. grep was grep. you didn't think i'd leave you hanging https://i.imgur.com/Vs5AAT7.png https://i.imgur.com/Vs5AAT7.png some other non-statistics from that day: 15GB sorted password list, newline delimited, UTF-8 from spinningrust drive 64 seconds (~234MB/s) to make a copy of the file. ag and rg took 3.2 seconds to search the copy. I'm actually hesitant to state that grep took 52 seconds... Thanks for replying, thanks for making me remember the great conversations we had around those topics a couple months ago, and thanks for creating ripgrep, it's my go-to for anything non-trivial!
- burntsushi 2y agoLove it! That's awesome. Thank you for replying. :-) I've occasionally wanted to put the subtitles from all of my Simpson episodes into an easily searchable format. What do you use to extract subtitles?
- genewitch 2y agoOh i've never extracted, i use openai-whisper for long content and whisper-diarization for shorter content (<8 minutes or so). As i suspected, ffmpeg claims to handle it: https://trac.ffmpeg.org/wiki/ExtractSubtitles https://trac.ffmpeg.org/wiki/ExtractSubtitles with a note that probably should be used with `-c copy` to ensure a 1:1 copy of the subtitles. also when i get stuff from a website with yt-dlp for archival i use ```pwsh $userInput = Read-Host -Prompt '480 video download script enter URL' Write-Output "URL:`t`t$userInput" yt-dlp.exe ` -f 'bestvideo[height<=480]+bestaudio/best[height<=480]' ` --write-auto-subs --write-subs ` --fragment-retries infinite ` $userInput ```