Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Language Driven Organ-Lesion Predictions in a Large Scale Toxicological Experiment
Learn how general-purpose language models predict compound toxicity by treating experimental factors as text, achieving accurate results without specialized models.
Predicting compound toxicity relies predominantly on representing molecules at various levels of
detail, with evaluation based on extensive experimental studies. These studies test a vast array of
factors to statistically estimate the potentially harmful outcomes of treatment conditions. This research
analyzes toxicology experiments from a new perspective. Rather than studying experimental factors
separately or developing specialized deep learning models, this approach interprets experimental
factors as strings and uses general-purpose pretrained language models for surprisingly accurate
predictions. Textual descriptors are either projected into a dense vector space with embedding models
and then integrated into a nested cross-validation pipeline, or given as input to state-of-the-art Large
Language Models (LLMs) that directly attempt zero-shot classification using schema-constrained
generation. The experimental validation uses data from the repeated-dose Open TG-GATEs dataset,
which exposes the same rat clone to 142 compounds over four different time periods at three levels
of dosage and contains histopathology annotations for both kidney and liver lesions.
Compose Email
Loading recent emails...