16 ms·
Show HN: BBC “In Our Time”, categorised by Dewey Decimal, heavy lifting by GPT
I'm a big fan of the BBC podcast In Our Time -- and (like most people) I've been playing with the OpenAI APIs.
In Our Time has almost 1,000 episodes on everything from Cleopatra to the evolution of teeth to plasma physics, all still available, so it's my starting point to learn about most topics. But it's not well organised.
So here are the episodes sorted by library code. It's fun to explore.
Web scraping is usually pretty tedious, but I found that I could send the minimised HTML to GPT-3 and get (almost) perfect JSON back: the prompt includes the Typescript definition.
At the same time I asked for a Dewey classification... and it worked. So I replaced a few days of fiddly work with 3 cents per inference and an overnight data run.
My takeaway is that I'll be using LLMs as function call way more in the future. This isn't "generative" AI, more "programmatic" AI perhaps?
So I'm interested in what temperature=0 LLM usage looks like (you want it to be pretty deterministic), at scale, and what a language that treats that as a first-class concept might look like.
- valgaze 4y agoThis is a really interesting use-case Applying "transformations" or classifying data in this way without having to setup a lot of detail-work seems like a real labor-saver/multiplier
- lerchmo 4y agoMy company has gone years wanting our product catalog to have structured data around our products but not going through the tedium of extracting it all. about an hour of prompt tweaking and it can pull, normalize, summarize and output valid json from all of our products. Basically pulling it out of a big unstructured html blob.
- genmon 4y agoI've started thinking about LLMs as a "universal coupling", if that makes sense? It's wild to be able to conceive of APIs to plain text, and natural language queries on structured APIs, but that's what we've got. My mind was really opened by Nat Friedman's work in GPT for browser automation: https://github.com/nat/natbot https://github.com/nat/natbot And of course using langchain/ReACT. So different from ChatGPT and (imo) way more intriguing. mentioned in this blog post: https://interconnected.org/home/2023/02/07/braggoscope https://interconnected.org/home/2023/02/07/braggoscope
- specproc 4y agoI love In Our Time, a real BBC gem. I've been meaning to pull all the audio for a while and this has inspired me. A fun thing to do would be to pass through Whisper, a great corpus to play with.
- genmon 4y agoI've been considering this, but my assumption is that it would be tripped up by the specialist words. I wonder... is there a way to "prime" Whisper (e.g. with the embedding of the episode synopsis) so that it "listens out" for words related to a particular topic? I haven't looking but this would be neat!
- taberiand 4y agoI haven't tried, but the Open AI docs mention priming on the whisper model being available prompt string Optional An optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
- coder543 4y agoIn my experience, Whisper does a great job even with specialized terminology. It won't catch everything, but I think it will exceed your expectations. One of the hardest things about Whisper is choosing which model to use; they offer a variety of sizes, and sometimes the smaller ones do better than the larger ones. It's worth trying a few different models and deciding what is best for each particular application. I will also say that I've personally been unimpressed with the new "large-v2" model, even though it supposedly scores better. The original "large-v1" model seems to work better than the "large-v2" model in the audio clips I've been testing Whisper against, but results will vary. In general, I find I'm really happy with what "small.en" and "medium.en" will emit, and they're much faster than the large models. (The ".en" models are specialized to English, and usually perform better for strictly English input, whereas the non-".en" models are trained on multiple languages.)
- haolez 4y agoWait. Is temperature=0 "pretty deterministic", or is it deterministic?
- LeoPanthera 4y agoIt’s my understanding that 0 is completely deterministic, except for when the model is updated, which does happen occasionally.
- valgaze 4y ago@haolez I just tested w/ davinci-003 w/ temperature to 0 Prompt: https://imgur.com/YtQ4fbf https://imgur.com/YtQ4fbf -- Reveal the question marks in an interesting way: The dog goes ????????????? -- With temperature == 0, it consistently ("pretty deterministically"?) generated "woof!" ex. The dog goes woof!
- gwillen 4y agoTheir docs say (somewhat recently updated, I think) that even 0 is not perfectly deterministic in all cases (though it's very close). Some people had previously observed this and speculated that it was some kind of floating point roundoff issue, when two outputs have almost identical scores.
- Jensson 4y agoGPU's often do floating point computations in nondeterministic order for performance reasons, so you would get small differences even for exactly the same order. It is probably that.
- genmon 4y agoIn the strict definition, it's deterministic: you get the same response for the same prompt, each time (given the exact same model). But the prompt is parameterised. The bulk of the prompt is requesting a list of guests and speakers to be extracted, and the episode synopsis is appended as the "parameter". And I've noticed that the variation of the parameter changes what the overall prompt returns... so it might start being less reliable at responding with valid JSON, for example. So it's instance-deterministic but, across a range of parameters, class-fuzzy, if that makes sense?
- PaulHoule 4y agoWhat's the prompt?
- valgaze 4y agoFrom what I understand: - Typescript definition (turn raw html markup >> well-formed JSON) - Dewey Decimal classification/score So probably combination of those two
- genmon 4y agoOne of the prompts is Extract the description and a list of guests from the supplied episode notes from a podcast. Also provide a Dewey Decimal Classification code and label for the description Return valid JSON conforming to the following Typescript type definition: { "description": string, "guests": {"name": string, "affiliation": string | null}[] "dewey_decimal": {"code": string, "label": string}, } Episode synopsis (Markdown): {notes} Valid JSON: (And the completion tends to be JSON, but not always.)
- lerchmo 4y agoI like the typescript definition, rather than example json that I normally use.
- genmon 4y agoCredit where it's due: I was working with structured data as JSON for the completion, and the Typescript definition hugely increased reliability. I took that from helpful advice (on Twitter) from Noah Brier who afiak came up with the approach: https://brxnd.substack.com/p/the-prompt-to-rule-all-prompts-brxnd https://brxnd.substack.com/p/the-prompt-to-rule-all-prompts-...
- noahbrier 4y agoOne interesting side effect of the Typescript interface approach is that it doesn't tokenize well: https://twitter.com/m1guelpf/status/1630015536632569857 https://twitter.com/m1guelpf/status/1630015536632569857
- dvt 4y agoThe BBC offers pretty exhaustive RSS feeds (https://podcasts.files.bbci.co.uk/b006qykl.rss https://podcasts.files.bbci.co.uk/b006qykl.rss) so I'm not exactly sure what ChatGPT even did here (Maybe the Dewey classification? Which is of dubious usefulness).
- genmon 4y agoThere's more on the About page but in summary: extracts the synopsis, guests (name, affiliation), and reading list (title, author, publisher, year) as structured data. Not massively hard with a bit of web scraping, but tedious and results in brittle code -- this took 20 minutes to write the prompt plus 3 cents per episode. https://genmon.github.io/braggoscope/about https://genmon.github.io/braggoscope/about (GPT-3 not ChatGPT for the model.)
- malborodog 4y agoHave you considered making the repo public so other folks (well, certainly me!) could extend and adapt the idea to other use cases? It's just very exciting work and it would be fun to see the code.
- sorokod 4y agoFor comparison two links for the (currently) last episode, first from BBC, the second from genmon https://www.bbc.co.uk/programmes/m001jkzg https://www.bbc.co.uk/programmes/m001jkzg https://genmon.github.io/braggoscope/2023/03/02/megaliths.html https://genmon.github.io/braggoscope/2023/03/02/megaliths.ht...
- aardvark179 4y agoI see several very dubious classifications. Shouldn’t the great stink be under civil engineering rather than agriculture, and why is Plato’s Atlantis under computer science?
- genmon 4y agoNow that's an interesting regression! I don't remember seeing it there before. (The worst I've noticed before has been Lawrence of Arabia under History of the Ancient World. Very much 20th century really.) Several other classifications are arguable -- which I think shows one of the limitations of this technique: it's not possible to iterate + improve. So instead I've been wondering about using the embeddings of each episode synopsis, and comparing to the embeddings of Dewey subdivisions. I should be able to tune the results better that way. There's also a technique from Google called CAVs (Concept Activation Vectors) that I'm intrigued about trying -- would love to hear if anybody has experience using this https://arxiv.org/abs/1711.11279 https://arxiv.org/abs/1711.11279
- theodorewiles 4y agoFYI you can get dewey decimals programmatically by hacking the OCLC API: http://classify.oclc.org/classify2/api_docs/classify.html http://classify.oclc.org/classify2/api_docs/classify.html this says you need an API key but I think I found some way to call this without one… You might be able to improve classification by either incorporating the dewey decimals of the books mentioned on a podcast or fine-tuning a model based on known book titles (or maybe there are book summaries somewhere) to known dewey decimals from OCLC.
- once_inc 4y agoThere's also an episode about aliens under 000.
- jgrahamc 4y agoWhat a beautiful thing.
- genmon 4y agoSome unlinked features... If you put the Dewey division in the URL, the directory auto-opens. e.g. here are episodes about prehistoric life (my current jumping-off point) https://genmon.github.io/braggoscope/directory#560 https://genmon.github.io/braggoscope/directory#560 There's a visual map of episodes. After principal component analysis of the episode embedding vectors, these are the most significant two components as the x,y https://genmon.github.io/braggoscope/map.html https://genmon.github.io/braggoscope/map.html (it's not super useful tbh -- e.g. the Manhattan Project and the Cambrian Explosion have the same x,y... presumably because they are both about explosions?) Many episodes have a reading list, and these are all linked to Google Books (so you can purchase/check out from a library), e.g. this episode page https://genmon.github.io/braggoscope/2022/10/20/the-fishtetrapod-transition.html https://genmon.github.io/braggoscope/2022/10/20/the-fishtetr... There are ~4,600 books, and I have ~88% coverage on getting a Google Books page from the original data. Any ideas about what to do with this big list of academic-recommended books v welcome!
- p0pcult 4y agoLove the visual map. What does color mean? Any way to do a 3rd PC, and put the visualization in a cube one can toy around with?
- genmon 4y agoColour is the 3rd component -- I wanted to see the difference between overlapping episodes. As for the 3D plot... here you go! https://interconnected.org/more/2023/03/in_our_time-PCA-3D-plot.html https://interconnected.org/more/2023/03/in_our_time-PCA-3D-p... Basic PCA + Plotly is actually in OpenAI's official Python library (in `embedding_utils`) -- this plot is just the output from that.
- p0pcult 4y ago:D this made my day.
- CrypticShift 4y agoExcellent! I'd love to see a script that sends a list of "descriptions" (1-100 words) to ChatGPT and directly gives you back a ready-made (embedding vectors closeness) map in a (textual) graph/chart format (like your above map or your plot https://interconnected.org/more/2023/02/in_our_time-PCA-plot.html https://interconnected.org/more/2023/02/in_our_time-PCA-plot...)
- boringg 4y agoAlso a big fan of the show -- cool project!
- astroalex 4y agoI love this project!! Ever since my partner and I discovered In Our Time a few years back, it’s been our go-to podcast to listen to together. Part of the allure is that the archive is so vast, but that makes it hard to browse. My partner made her own archive of In Our Time here, if you’re interested: https://shelby.cool/melvyn/ https://shelby.cool/melvyn/ She used Wikipedia to find and categorize each episode. I also really like that she indexed episodes by guest, too. Certain guests are REALLY good and have been on many episodes. Super excited to see someone else make an archive; we’ll definitely be exploring yours!
- genmon 4y agoNo way! This is incredible. That h1 "Hello," <3 Is the tagging manual? It's really good.
- astroalex 4y agoIt's definitely not manual, here are the scripts she used to generate the archive: https://github.com/shelbywilson/melvyn/tree/main/scripts https://github.com/shelbywilson/melvyn/tree/main/scripts
- zeristor 4y agoThat's really impressive. I've had a mild idea for ages to have a sort of annotated In Our Time. Listening to the podcast on a webpage as the text rolls by links could appear to explain or give background to the item or person being discussed. SMIL is the multi-media mark up language. Generally if one thinks of something there's someone on the Internet who has already had that idea. Additional: I think the BBC is very careful about transciptions. They've sold a book of the transcripts of several episodes, but it would be a great way to go through a subject.
- bscphil 4y ago> She used Wikipedia to find and categorize each episode. This is a really clever use of an existing dataset. I clicked through before reading this and was stunned by how thorough the tag set was. Even more obscure things like "Alumni of Magdalen College, Oxford" have multiple episodes. I'm going to keep this in mind on future projects for sure.
- theodorewiles 4y agoI emailed you but did something similar and posted to Show HN a while back - https://weekend-collection.s3.amazonaws.com/Catalog+-+Feb+17.html https://weekend-collection.s3.amazonaws.com/Catalog+-+Feb+17...
- genmon 4y agoJust to make sure that others see this, I found your clustering technique super interesting and very effective, and something I want to try myself. Your technical writeup is here: https://weekendcollection.substack.com/p/technical-details https://weekendcollection.substack.com/p/technical-details (Recursive coarse clustering as opposed to one-shot fine clustering.)
- theodorewiles 4y agothanks!
- jccalhoun 4y agoIt is an interesting approach and I appreciate it. However, an alphabetical list would be more useful for me when I am interested in topics and might not know where they are classified.
- jack_riminton 4y agoIn Our Time has been a life-long companion, one of things that makes me proud of the BBC
- genmon 4y agoHave you read the New Yorker articles about it? On _In Our Time_ -> https://www.newyorker.com/culture/podcast-dept/escape-the-news-with-the-british-podcast-in-our-time-with-melvyn-bragg https://www.newyorker.com/culture/podcast-dept/escape-the-ne... Profile of Melvyn Bragg -> https://www.newyorker.com/culture/the-new-yorker-interview/the-education-of-melvyn-bragg https://www.newyorker.com/culture/the-new-yorker-interview/t... Both well worth your time.
- jack_riminton 4y agoThat NewYorker summarises it nicely. I can’t stand the over-produced stuff on NPR for example. In Our Time has no politics, no hook to get you listening for the next episode. It’s just a lovely small window into academia. Above all it makes things simple without dumbing it down
- gerdesj 4y ago"It’s just four intelligent people in a studio, discussing complex topics that are, as a friend of mine once said of Bragg’s openers, aggressively uncommercial." Nicely put by the New Yorker.
- EamonnMR 4y agoIn Our Time is one of the best podcasts out there. Are there any similar ones?
- open-source-ux 4y agoSimilar to In Our Time is another BBC radio programme called The Forum (from BBC World Service) which explores world history, culture and ideas. Every week an eclectic topic is discussed with three experts in a lively, informative and stimulating discussion. Highly recommended: https://www.bbc.co.uk/programmes/p004kln9/episodes/downloads https://www.bbc.co.uk/programmes/p004kln9/episodes/downloads
- maccaw 4y agoLove it - also a huge fan of In our Time. V2 would be: - Transcribe through Whisper - Semantic search
- DC-3 4y agoNice project OP. I also love In Our Time. Some favourite episodes off the top of my head: * Wilfred Owen - https://www.bbc.co.uk/programmes/m001df48 https://www.bbc.co.uk/programmes/m001df48 * The Evolution of Crocodiles - https://www.bbc.co.uk/programmes/m000zmhf https://www.bbc.co.uk/programmes/m000zmhf * The May Forth Movement - https://www.bbc.co.uk/programmes/m001282c https://www.bbc.co.uk/programmes/m001282c * The Valladolid Debate - https://www.bbc.co.uk/programmes/m000fgmw https://www.bbc.co.uk/programmes/m000fgmw * Gerard Manley Hopkins - https://www.bbc.co.uk/programmes/m0003clk https://www.bbc.co.uk/programmes/m0003clk * Henrik Ibsen - https://www.bbc.co.uk/programmes/b0b42q58 https://www.bbc.co.uk/programmes/b0b42q58 * Wuthering Heights - https://www.bbc.co.uk/programmes/b095ptt5 https://www.bbc.co.uk/programmes/b095ptt5 And finally, in which three mathematicians heroically attempt to explain asymptotic analysis to (septuagenarian novelist and cultural broadcaster) Melvyn: * P v NP - https://www.bbc.co.uk/programmes/b06mtms8 https://www.bbc.co.uk/programmes/b06mtms8
- genmon 4y agoGreat list! Some personal faves in return, in no particular order: - The evolution of teeth https://www.bbc.co.uk/programmes/m0003zbg https://www.bbc.co.uk/programmes/m0003zbg - The fish-tetrapod transition https://www.bbc.co.uk/programmes/m001d56q https://www.bbc.co.uk/programmes/m001d56q - The late Devonian extinction https://www.bbc.co.uk/programmes/m000sz7x https://www.bbc.co.uk/programmes/m000sz7x - The American West https://www.bbc.co.uk/programmes/p00548gg https://www.bbc.co.uk/programmes/p00548gg - Metamorphosis (Ovid) https://www.bbc.co.uk/programmes/p00546p6 https://www.bbc.co.uk/programmes/p00546p6 - Politeness https://www.bbc.co.uk/programmes/p004y29m https://www.bbc.co.uk/programmes/p004y29m - The Bronze Age collapse https://www.bbc.co.uk/programmes/b07fl5bh https://www.bbc.co.uk/programmes/b07fl5bh - Doggerland https://en.wikipedia.org/wiki/Doggerland https://en.wikipedia.org/wiki/Doggerland
- dendrite9 4y agoI really enjoyed the recent Superconductivity episode. Especially hearing Melvyn say 'Good god' half way through.
- martinpw 4y ago
- mattlondon 4y agoFinally, my interest in LLMs is piqued! Seems like everyone has been getting excited around the search or code-generation use cases ... or simply trying to make it say naughty things (boring, not interested, wake up in a few more years), but this is eye opening. The idea of this as a "universal coupler" is fascinating, and I think I agree with the author that we are probably standing at an early-90s-web moment with LLMs as a function call (the technology is kinda-there and mostly-works, and people are trying out a lot of ideas ... some work, some don't). My mind is racing. Thanks for the epiphany moment.
- ketzo 4y agoIt’s awesome for formatting and structuring. Copy and paste a bunch of styled JS components -> get back out a single CSS sheet Paste in a markdown document -> get out the same thing in HTML Fun stuff like that.
- chaxor 4y agoAh yes, you now too can have the wonders of pandoc - now on sale! Retail pandoc price of 20MB, now selling for only 3TB with the exclusive offer of ChatGPT!
- seagullriffic 4y agoI don't want to be mean, but this seems like the famous Dropbox/rsync comment. The value here is how easy it is, and the fact that a generalised model can take the place of a (well-engineered) specialised tool.
- v7n 4y agoBit of a tangent but I also predict an explosive demand for parametric AI-first design and simulation software. OpenSCAD works right now in the sense that the produced source is valid, but asking for even something well documented e.g. an AR-15 lower receiver does not impress.
- 4y ago
- simonw 4y agoI wanted to see which speakers had been on the most episodes. I went to https://genmon.github.io/braggoscope/guests https://genmon.github.io/braggoscope/guests and opened the Firefox DevTools console and ran this: guests = Array.from(document.querySelectorAll('ul a.text-blue-500.underline')).map(el => ({ name: el.innerText, count: parseInt(el.nextElementSibling.textContent.slice(1, -1), 10) })) Then this: guests.sort((a, b) => a.count < b.count) console.log(JSON.stringify(guests, null, 2)) The top few were: [ { "name": "Simon Schaffer", "count": 24 }, { "name": "Angie Hobbs", "count": 23 }, { "name": "Martin Palmer", "count": 22 }, { "name": "Steve Jones", "count": 21 }, { "name": "Paul Cartledge", "count": 20 }
- genmon 4y agoI think Angie Hobbs will take it! Looks like she is referred to with a few different names in the episode notes and I’ll need to merge them manually (will be 25 appearances) (btw thanks for datasette. I use SQLite as an intermediary db and datasette was invaluable for exploring and refining queries.)
- simonw 4y agoAny chance you might share the intermediary DB? Even just including the binary database file on GitHub somewhere would be neat, since I could play around it by doing https://lite.datasette.io/?url=URL-to-your-db-file-on-GitHub https://lite.datasette.io/?url=URL-to-your-db-file-on-GitHub
- I_complete_me 4y agoLots of great contributors to the show over the years. My favourite is Steve Jones for his huge intellect, humility, knowledge and listenabilty - yeah it's the right word for someone like him. I've listened to a lot of episodes (some many times over). I'll just mention The Migration of Birds as a sound listen. However, Melvyn is an "interesting" catalyst but still great to fall asleep to. Long may it continue. I believe there is sister program with a younger female presenter. Name escapes me...
- donut 4y agoThis is so cool! Thanks for sharing. Taxonomies are inherently limited -- I love this portion of the talk "Everything is Miscellaneous" about Melvil Dewey: https://www.youtube.com/watch?v=x3wOhXsjPYM&t=1206s https://www.youtube.com/watch?v=x3wOhXsjPYM&t=1206s Surely some things fit equally well in more than one category. Have you considered asking for best two or three categories, and placing the episodes in multiple locations? Or would that be too noisy?
- genmon 4y agoI found that Dewey was borderline acceptable — more categories seemed to degrade it, though admittedly I didn’t spend much time prompt tuning. I also tried tags and these weren’t reliable (not at a consistent “scale” from episode to episode). I suspect I’ll need a more mechanical approach, long term.
- vidro3 4y agoI always assumed the podcast was current with the show but it seems like it's about 4 weeks behind. e.g. I have Tycho Brahe on March 2 in my podcast app, rss shows it released Feb 2.
- cheschire 4y agoNo longer does Tony the Pony need to parse XHTML with regex. Who knew GPTs would be the actual solution?
- haunter 4y agoSomething like this for Essential Mix episodes with tracklists and genre would be mindblowing
- SomewhatLikely 4y agoI'm not sure this one about sewage is correctly classified. It's in agriculture, presumably because it mentions fertilizing fields. It probably fits better under engineering. https://genmon.github.io/braggoscope/2022/12/29/the-great-stink.html https://genmon.github.io/braggoscope/2022/12/29/the-great-st...
- viburnum 4y agoI love it! "In Our Time" is the first podcast I listened to, way back in 2004, and I still listen to every episode.
- arsome 4y agoTea under Economics and coffee under Technology > Home and family management... some... interesting decisions in here.
- ada1981 4y agoSo ChatGPT was used in prepping the data but isn’t used dynamically correct? You may be able to create a bot that chats with you and suggests episodes, etc. in real time. Or lets you ask questions TO an episode. And other things.
- nextaccountic 4y ago> Web scraping is usually pretty tedious, but I found that I could send the minimised HTML to GPT-3 and get (almost) perfect JSON back: the prompt includes the Typescript definition. Could you share the prompt? Or, if OP can't share, does anyone have ideas for a prompt to do something like this?
- kwerk 4y ago+1, and the OP mentions a wrapper to handle invalid JSON from GPT-3. I’d be interested in that too OP’s other write up: https://interconnected.org/home/2023/02/07/braggoscope https://interconnected.org/home/2023/02/07/braggoscope
- malborodog 4y agoThe prompt is probably simple, but the bigger challenge is that even a minified html of a typical web page would be more than the 4k gpt token limit
- genmon 4y agoI extract the main content div, which includes various other divs and assorted HTML cruft from 25 years of content management systems. Then convert that to Markdown, which GPT groks happily, and it preserves the right balance of discarding meaningless structure but preserving some semantics (italics, headings, etc). The best tool I've found for that process is aaronsw's html2text, amazing that it's still so valuable after all these years.
- malborodog 4y agoThanks for explaining — very helpful!
- genmon 4y agoShared in this comment https://news.ycombinator.com/item?id=35073824 https://news.ycombinator.com/item?id=35073824
- megablast 4y ago> Web scraping is usually pretty tedious, but I found that I could send the minimised HTML to GPT-3 and get (almost) perfect JSON back What does this even mean?? You collected all the minimised html snippets then fed them to GPT-3?? How does this save any time??
- totetsu 4y agoyesterday I had quite a lot of success giving it some json and asking for a python script to make an Anki deck from it. The programmatic use case is very magical
- lisasays 4y agoAt the same time I asked for a Dewey classification... and it worked. May I ask - how do you know exactly? Specifically - with what level of accuracy? And how is it assessed? At first glance it does seem to be highly accurate - except where it isn't: 650 Management and public relations (1) - Caxton and the Printing Press 18 Oct, 2012 Which is not to knock what you're doing - because overall the performance does seem to be quite impressive (especially considering the effort invested). Still - there's that "last mile". So it would be nice to now how the performance compares to that other topic modeling approaches. And to what extent LLMs actually provide a boost.
- SnoJohn 4y agoThis is very cool. Would love to try it on The Rest is History podcast.
- deleted 4y ago[deleted]
- dirtyid 4y agoThis is great.
- lysecret 4y agoYea LLM for data Extraction will be huuuuuge. I had so many Data Projects in the past which boiled down to: Damn the data is too unstructured.
- coob 4y ago90 episodes on the History of Europe and 3 on the History of Africa. Blimey
- kzrdude 4y agoThere's a lot of Egypt episodes missing from that category, for example.
- ImaCake 4y agoTo be fair, the podcast is produced by Europeans. Still, a British history podcast should spend more time recogning with its colonial history, and they have done at least one very very good episode on it.
- etewiah 4y agoWell spotted!! My degree was in African Studies and I remember one of the first lectures talked about a famous Oxford historian (Hugh Trevor Roper?) who said something like - there is no history in Africa, only the wild gyrations of savage tribes..... Made me very sad.
- davidy123 4y agoThe Dewey system is famously stuck in its time and culture. Look at religion, there are multiple top level cateogries for Christian religion related things, then there's "Other religions" for everything else. At this point, it's embarassing and insulting.
- thom 4y agoYeah, I've found that ChatGPT is excellent at filling out XML templates too: Fill out the following XML template to describe a room in a fantasy RPG dungeon: <room width="" length=""><furniture><object name="" description=""/></furniture><npcs><npc name="" description=""/></npcs><items><item name="" description=""/></items></room> This yields, for example: <room width="10" length="15"> <furniture> <object name="torch" description="A wooden torch mounted on the wall, casting flickering light around the room."/> <object name="table" description="A large wooden table in the center of the room, covered with maps, scrolls, and dusty tomes."/> </furniture> <npcs> <npc name="Goblin" description="A small, green-skinned creature with a wicked grin and a rusty dagger in hand."/> <npc name="Orc" description="A towering figure with rippling muscles, wielding a massive axe and snarling menacingly."/> </npcs> <items> <item name="Potion of Healing" description="A small vial filled with a glowing green liquid, said to restore vitality when consumed."/> <item name="Sword of Flame" description="A gleaming sword with a blade wreathed in flames, capable of burning foes to a crisp."/> </items> </room> So you get the best of both worlds, really clever procedural generation of content but also a tight structure - I've yet to see a syntax error out of it.
- DeathArrow 4y agoBut isn't classification a simple problem solved long time ago with simple ML techniques such as SVMs and neural networks? I remember having done a project for a ML class in Uni and I used a SVM. It took movies and they were classified in genres. Training data was IMDB comments.
- thewarrior 4y agoHow long did it take you ? This tech turns days of work into hours or minutes of work.
- DeathArrow 4y agoA week or something.
- gadders 4y agoThis is pretty cool. If you like In Our Time, I would also suggest the BBC World Service equivalent, The Forum. https://www.bbc.co.uk/sounds/brand/p004kln9 https://www.bbc.co.uk/sounds/brand/p004kln9 It's also a shame the BBC delays the In Our Time podcast by a month unless you use their shite BBC Sounds app (which I think only exists to make the Spotify app look good in comparison).
- corobo 4y agoHere come the real AI use cases! Nice work! AI generated content is the initial "ooh, money" wave that will be the future's basic spam to deal with. Difference between AI and web3 is that web3 didn't have a real use case to follow the gold rush phase :x
- Southworth 4y agoHey Matt! Great to see this on HN, we’ve been chatting away about this on a private slack group! Super cool work as ever. Stay awesome.
- jgtrosh 4y agoSome episodes that have subtitles are titled incorrectly (only using the subtitle). For example, this episode [1] is the first of a pair of episodes called The Written World, but it only shown as Episode 1. [1]: https://genmon.github.io/braggoscope/2012/01/02/episode-1.html https://genmon.github.io/braggoscope/2012/01/02/episode-1.ht...
- ourmandave 4y agoWeb scraping is usually pretty tedious, but I found that I could send the minimised HTML to GPT-3 and get (almost) perfect JSON back: the prompt includes the Typescript definition. This is the part that sounds amazing to me. I haven't done web scraping since forever, but it was tag matching hell even on a consistent data source. And that it hands you back JSON is just unfair.
- mymythisisthis 4y agoShould be arranged in chronological order. That way you can listen to the entire history of the world.
- Janymos 4y agoI wonder what the standard approach is when the LLM does not return valid JSON data? Do you skip the input data all together, or use the parsing error to generate a valid JSON?
- genmon 4y agoThe general rule is: be minimally liberal in what I receive :) A parse error kicks off a recovery process where, in theory, we could run any number of rules. In practice the only problems are unescaped quotes in strings, or mismatched quotes (start with `"` and terminate with `'`)
- OOPMan 4y agoHow perfect is almost perfect? How do you correct errors?
- lysecret 4y agoI think it’s crazy to think about that we scraped all this data from the web to train this crazy LLM to be able to interpret the web and scrape more data from it.
- joshspankit 4y agoThis is amazing, and exactly the kind of thing I’ve been hoping for to help us find the gems in the endless stream of excellent content. What would it take to go deeper on this and narrow down to, say, single-sentence intervals? For example finding everything about a particular character in the Ramayana, or every statement about NPR itself?
- rafazinho 4y ago[dead]
- gnutmal 4y agoThis is amazing, thanks for sharing!