Inspiration

Taking a promising idea from the lab to an approved medicine can take roughly 10 to 15 years. And along the way, researchers may have to make and test hundreds and thousands of chemical compounds just to find one worth moving forward. This made us ask ourselves how this process could be optimized so that people did not have to wait over a decade for life saving medical discoveries and led to us building Estrogen Receptor blocker. We decided to focus on breast cancer because around 70 to 80 percent of breast cancers in women are ER-positive. In these cancers, estrogen can bind to the estrogen receptor and help drive tumour growth, which is why blocking this pathway is such an important treatment strategy.

What it does

Scientists might already have a molecule that works reasonably well, too toxic, too lipophilic, or one that they want to retain the useful parts of the molecule while increasing its predicted activity against ERα. They can feed the system with a molecule of interest, and from there, a partial strand can be expanded to possible completions of a valid molecule. It can also calculates properties of a complete molecule, including molecular weight, toxicity, LogP, hydrogen-bond donors, hydrogen-bond acceptors, and synthetic-accessibility score. The goal of this tool is to provide a tool to work as a research prototype, giving a smarter starting point which will help them eliminate weaker candidates earlier, and focus their time and resources on the most promising molecules.

How we built it

The partial molecule generation is done through a transformer model ChemGPT, which was trained on chemical sequences to generate possible completions. Results are passed through RDKit which checks whether the structure is chemically valid, canonicalizes it, removes duplicates, and calculates properties. We use a random forest classier to train with datasets ChEMBL and ClinTox, to predict the likelihood a molecule is to interact strongly with the estrogen receptor, and the likelihood that the molecule has high toxicity, respectively. The whole process is visualized on a web-based interactive platform, built with Python and JavaScript.

Challenges we ran into

It was not easy to acquire the datasets we needed, and to train the classifier on our computers without issues. Obtaining a viable artifact set took multiple iterations, involving many changes to the original plan.

Accomplishments that we're proud of

Our two classifiers (on ChEMBL and ClinTox) are surprisingly accurate when tested against the dataset, consistently reaching over 87% on accuracy and precision. We are also proud of the outputs that our model can predict, and how potentially useful the project can be in drug discovery.

What we learned

We learned the importance of finding reliable datasets, cleaning and preparing data, and making sure generated molecules are chemically valid. We also learned how different machine learning models can be combined with existing tools like ChemGPT and RDKit to solve a larger problem. In the end, we learned that building a useful research tool requires constant testing and iteration, and that even when an approach does not work as expected, it can help us improve the overall design.

What's next for Estrogen Receptor Blocker

Since our ultimate goal is to optimize drug development process, we would like to target health issues beyond breast cancer. Using AI statistics as a tool in healthcare to allow faster advancement in life saving medical discoveries is what we envision.

Built With

Share this project:

Updates

Submission history