5 ms·
This match was broadcast by ESPN on both ESPN2 and ESPN+. In that case, they are presumably paying two knowledgeable commentators to talk about the match before
by jetrink 2y ago
This match was broadcast by ESPN on both ESPN2 and ESPN+. In that case, they are presumably paying two knowledgeable commentators to talk about the match before, during, and after it happens, adding context and describing the important events. Are they not providing a transcript of that to the writerbot? That seems like real low-hanging fruit, especially since the commentary is already live captioned.
- batesy 2y agoThis is what I was thinking too... Still early days I guess lol
- nolok 2y agoIt's a company that see itself as "media / journalism", and having consulted for a few of those I've always been amazed that their tech teams is most often isolated from the content team with very low access to said media. It's very different from what you would expect in tech (access to everything), or just common sense in general. Note that I have no knowledge whatsoever of how ESPN work, I'm inferring from what I've seen elsewhere.
- btown 2y agoSports have so much structured data, and such a high bar for describing it accurately (especially for a brand like ESPN), that there are significant risks to the hallucinations that might develop from a multi-hour transcript being fed into an LLM, especially with commentators excited about potential goals and other events that don't end up happening. On the other hand, the rather simple task of "here's a set of goals, their times, who made them, who assisted... turn that into prose" could even be done without LLMs with a deterministic algorithm, and may very well have been in this case. Some of the grammar issues in the OP feel very pre-LLM in nature, like a combination of substitution rules gone awry. Now, could you create a system that repeatedly interrogates the statements made by a first pass of an LLM on summarizing a long transcript, and comparing those results against structured data you know for accuracy? Would this lead to richer content and accessible error rates relative to the simpler approach? Would this be the type of thing that the best machine learning engineers in the world could probably prototype over a hackathon? The answer is very possibly yes to all three of these. But it's far from low-hanging fruit for any sizable, risk-averse organization. It's very difficult to fight against "the thing we have is imperfect, but at least it never gets the facts wrong."
- btown 2y agos/accessible/acceptable/ - guess I should have run my comment through the kind of check-for-typos LLM step that I described above!
- jetrink 2y agoI got nerd-snipped by this. I transcribed the match using the whisper-small.en and then asked ChatGPT to create a summary using a neutral prompt: Here is a transcript of a soccer match. In the style of an experienced professional sports reporter, please write a 200 word article about the match. Its summary starts, "In her final professional match, Alex Morgan delivered a performance filled with emotion and resilience, though her San Diego Wave fell short in a 3-1 loss to North Carolina Courage. The game at Snapdragon Stadium in San Diego was more than just a contest; it was a tribute to one of soccer’s most iconic figures." It did get the score wrong. Here's the rest: 1. https://chatgpt.com/share/de8c60d1-69ab-4291-99dc-d4d95af3d324 https://chatgpt.com/share/de8c60d1-69ab-4291-99dc-d4d95af3d3...
- regretaverse 2y agoOpenAI transcription & LLM is likely more costly than what they're willing to splurge on this. Care to link soccer.txt, so I can try with Llama 3 / 3.1?
- whimsicalism 2y agotoo costly? it’s probably a few cents total
- brewdad 2y agoAnd it got the most basic detail, the score of the game, wrong. Even at the cost of only a few cents it's worthless.
- whimsicalism 2y agoseems like this is a solvable problem with just a bit more engineering effort, but yeah... totally worthless
- acchow 2y agoYou should just need a better prompt. I think everyone would benefit from using a standardized prompt which asks the model to think through its work between `<thought>` tags before writing its response, and also reflecting on the response between `<reflection>` tags, and then outputting the final response afterwards