5 ms·
Here is a very real example where 32k tokens is not enough: predicting discharge disposition for a long stay inpatient. A 30-day hospital stay can produce over
by pilotneko 3y ago
Here is a very real example where 32k tokens is not enough: predicting discharge disposition for a long stay inpatient. A 30-day hospital stay can produce over 1,000 clinical notes.
- jerrygenser 3y agoWhen you say prediction my discharge disposition do you mean extracting from the document a statement of what the discharge disposition was? Or do you mean based on m extracted clinical characteristics, predicting what it should be? For extracting what it was if it's in the document, you can use the indexing approach described in the article. For predicting what it should be, you'd probably be better off using llm as an extraction mechanism and then using structured data models to classify discharge disposition
- pilotneko 3y agoThe latter, not the former. The former is extraction, not prediction. A lot of factors go into predicting the post-acute care route that results in the best outcome. You are correct that it would be possible to create features using an LLM, but that is a very difficult problem compared to simply treating the problem as a sequence classification.
- jerrygenser 3y agoAgreed that there are better approaches than LLm to solve this problem There are deep learning and sequence models that are focused on this task using extracted structured info. Wouldn't you be better off using llm or any extraction + linking model to identify medical concepts and then a model that only understands medical concepts in order to do predict the next medical concepts.
- lmeyerov 3y agoYou can do layered retrieval and summarization/embeddings to fit within the context bounds. Conceptually it's not far from what a clinician would do anyway. By switching from a naive one shot prompt with big context to interactive (multiple calls) and summaries, no need to retrain But yeah my money is on context windows getting bigger and frameworks smoothing out how to do the above automatically, So it feels like a point in time optimization right now That will be tech debt by the end of the year
- avereveard 3y agoIs this a joke? Medical data should stay the hell away from gpt servers
- pilotneko 3y agoOpenAI will sign a business agreement and run an instance with isolated servers, for HIPAA compliance. That being said, you can run LLMs on-prem.
- avereveard 3y agocitation needed. specifically regarding gpt4 and not the other models, since that was your claim. > That being said, you can run LLMs on-prem. how is this relevant at all? it's just dodging the question, since the claim was about gpt4 32k tokens and not the crop of generic models that we have now availabe on prem (and which don't support 32k token anyway)
- famouswaffles 3y agohttps://arstechnica.com/information-technology/2023/04/gpt-4-will-hunt-for-trends-in-medical-records-thanks-to-microsoft-and-epic/ https://arstechnica.com/information-technology/2023/04/gpt-4...
- avereveard 3y agofrom that source we find descriptions of two features "enhancements to automatically draft message responses" "Another solution will bring natural language queries to SlicerDicer" neither of which needs 32k token nor to see the patient records also, foot note 1 "users of Azure and Azure OpenAI Service are responsible for ensuring the regulatory compliance of their use" there is zero claim of actual compliances of these services for handling sensitive or regulated data. now I understand the enthusiasm, but at least don't waste people time with sources that are, at best, tangential, and don't provide any substance to the discussion
- famouswaffles 3y ago