3 ms·
Hi all, author here! My submission fell of the new page after 5 minutes sometime this morning and really glad to see it get re-posted. AMA!
by binarymax 5y ago
Hi all, author here! My submission fell of the new page after 5 minutes sometime this morning and really glad to see it get re-posted.
AMA!
- bluetwo 5y agoI have worked with eCFR and have thought about some ideas for processing large sections and how it might be useful. What are your larger plans here? Do you have an interface in mind for searching or organizing things?
- binarymax 5y agoOh yeah I could tinker forever, it's an amazing dataset that I think needs more attention from the ML community. Glad to see the working team at https://www.ecfr.gov/ https://www.ecfr.gov/ finally making their search better, as Cornell Law has been the defacto go to forever (for me at least). I think an amazing eCFR search experiment would be transformer vectors in a graph, using the hierarchy, citations, and references as edges to (sub)paragraph and section nodes - perhaps even using a modified HNSW somehow. The graph that exists there now isn't leveraged enough. Per this dataset itself, I already output to Vespa formatted JSON (as noted in https://github.com/maxdotio/ecfr-prepare https://github.com/maxdotio/ecfr-prepare )...and the resulting vectors from the inference get appended to the original JSON doc as a field. I have a Vespa schema hat I need to upload (that doesnt include the vector field yet but can be added using the Vespa vector search walkthroughs). It's been a busy day but I'll quickly try to find a place to put it for now :) --EDIT-- Pushed the schema to the above repo, and some bash. You'll need Docker and to follow the Vespa MSMARCO instructions first at https://docs.vespa.ai/en/tutorials/text-search-semantic.html https://docs.vespa.ai/en/tutorials/text-search-semantic.html to get used to the engine.
- bluetwo 5y agoYes, I was happy to see the modern changes, I agree Cornell Law had done a better job, although I think a lot of people use Google as the search tool and then link to their prefered site, since they are always the first two. My experience has been with 14 CFR and 21 CFR. I would love to see any tool you come up with in the future and would be happy to give you feedback.
- Gollapalli 5y agoI think vectorizing and eventually algorithmizing federal regulations is important work and will be important for outcome driven federal policy, and I’m glad that you’re doing it.
- binarymax 5y agoThanks! I hope newcomers see the importance, beauty, and complexity of the dataset and run with it as well. The more interested the better. "Augmented Federal Register" would be of real help for better crafting final rules as well - which is another area I'm looking into - as FR is truly organic.
- KarlKemp 5y agoIt's completely impossible without strong AI. As but one example: contracts often include the phrase "a reasonable effort". Here's a definition: "Reasonable Efforts means, with respect to a given goal, the efforts that a reasonable person in the position of the promisor would use so as to achieve that goal as expeditiously as possible" Try defining that in an algorithm!
- binarymax 5y agoI don't think full-automation would even be a goal - not when crafting laws for humans! But perhaps better tools can be built to make the process more efficient, less biased, and less redundant.
- turnersr 5y agoAwesome work!!!! curious how you handle paragraphs and niche language like federal regulations. What are your favorite ways to do sentence and paragraph embeddedings and is there a framework you like where you can tune to custom data? Do you find fine tuning your embedding model helpful?
- binarymax 5y agoThanks! The post doesn’t cover fine tuning of the model which would be absolutely necessary (but out of scope for the post). Nils Reimers (the author of SBERT) has been on a speaking circuit covering Generative Pseudo Labelling to handle the vocabulary gap of new domains that a pretrained sbert model hasn’t seen yet. https://youtu.be/qzQPbIcQu9Q https://youtu.be/qzQPbIcQu9Q
- aghilmort 5y agogreat project! definitely curious about this for lots of reasons; can add as resource to our legal topic filter list at Breeze search whenever you go live; may also be some collab opportunities