Preprint
Machine Learning

OPERA models for predicting physicochemical properties and environmental fate endpoints

Kamel Mansouri(Oak Ridge Associated Universities), Chris Grulke(Environmental Protection Agency), Richard Judson(Environmental Protection Agency), Antony Williams(Environmental Protection Agency)
March 8, 2018Journal of Cheminformatics614 citations

614

Citations

30

Influential Citations

Journal of Cheminformatics

Venue

2018

Year

Abstract

The collection of chemical structure information and associated experimental data for quantitative structure–activity/property relationship (QSAR/QSPR) modeling is facilitated by an increasing number of public databases containing large amounts of useful data. However, the performance of QSAR models highly depends on the quality of the data and modeling methodology used. This study aims to develop robust QSAR/QSPR models for chemical properties of environmental interest that can be used for regulatory purposes. This study primarily uses data from the publicly available PHYSPROP database consisting of a set of 13 common physicochemical and environmental fate properties. These datasets have undergone extensive curation using an automated workflow to select only high-quality data, and the chemical structures were standardized prior to calculation of the molecular descriptors. The modeling procedure was developed based on the five Organization for Economic Cooperation and Development (OECD) principles for QSAR models. A weighted k-nearest neighbor approach was adopted using a minimum number of required descriptors calculated using PaDEL, an open-source software. The genetic algorithms selected only the most pertinent and mechanistically interpretable descriptors (2–15, with an average of 11 descriptors). The sizes of the modeled datasets varied from 150 chemicals for biodegradability half-life to 14,050 chemicals for logP, with an average of 3222 chemicals across all endpoints. The optimal models were built on randomly selected training sets (75%) and validated using fivefold cross-validation (CV) and test sets (25%). The CV Q2 of the models varied from 0.72 to 0.95, with an average of 0.86 and an R2 test value from 0.71 to 0.96, with an average of 0.82. Modeling and performance details are described in QSAR model reporting format and were validated by the European Commission’s Joint Research Center to be OECD compliant. All models are freely available as an open-source, command-line application called OPEn structure–activity/property Relationship App (OPERA). OPERA models were applied to more than 750,000 chemicals to produce freely available predicted data on the U.S. Environmental Protection Agency’s CompTox Chemistry Dashboard.

Analysis

Why This Paper Matters

This paper addresses a critical bottleneck in environmental chemistry and regulatory science: the lack of high-quality, reproducible quantitative structure-activity/property relationship (QSAR/QSPR) models for physicochemical properties and environmental fate endpoints. With over 614 citations, it has become a foundational resource for computational toxicology and green chemistry. The work is particularly significant because it adheres to the five OECD principles for QSAR validation, ensuring that the models are not only statistically robust but also mechanistically interpretable and suitable for regulatory decision-making. By making all models open-source via the OPERA command-line application and applying them to more than 750,000 chemicals on the EPA CompTox Chemistry Dashboard, the authors have democratized access to reliable property predictions, enabling researchers and regulators to fill data gaps without expensive experiments.

Technical Contributions

  • Data Curation Workflow: Developed an automated pipeline to curate the PHYSPROP database, standardizing chemical structures and selecting only high-quality experimental data for 13 endpoints (e.g., logP, boiling point, biodegradability half-life).
  • Weighted k-Nearest Neighbor (kNN) Modeling: Adopted a weighted kNN approach that balances local structure-activity relationships, using a minimal set of 2–15 descriptors (average 11) selected by genetic algorithms for mechanistic interpretability.
  • OECD Compliance: Models were validated by the European Commission's Joint Research Center to meet all five OECD principles, including defined endpoint, unambiguous algorithm, applicability domain, goodness-of-fit/robustness/predictivity, and mechanistic interpretation.
  • Open-Source Implementation: Released OPERA as a free, command-line tool using PaDEL descriptors, enabling reproducible predictions and integration into larger workflows.

Results

The models achieved strong predictive performance across 13 endpoints. Cross-validation Q² values ranged from 0.72 to 0.95 (average 0.86), and test set R² ranged from 0.71 to 0.96 (average 0.82). Dataset sizes varied widely: from 150 chemicals for biodegradability half-life to 14,050 for logP, with an average of 3,222 chemicals per endpoint. These metrics indicate high robustness and generalizability, especially given the diversity of endpoints (e.g., vapor pressure, soil sorption, aerobic biodegradation). The models were applied to over 750,000 chemicals, providing predicted data freely available on the EPA CompTox Chemistry Dashboard.

Significance

OPERA has had a transformative impact on computational environmental chemistry by providing a validated, open-source framework for predicting key properties without experimental testing. Its adherence to OECD principles makes it suitable for regulatory submissions, reducing reliance on animal testing and costly experiments. The integration with the EPA CompTox Dashboard has made these predictions accessible to a broad community, accelerating risk assessment and chemical prioritization. For AI practitioners, the work demonstrates how classical machine learning (kNN) combined with rigorous data curation and descriptor selection can achieve high performance and regulatory acceptance, offering a template for building trustworthy predictive models in scientific domains.