ArticleslgStudy

chemistry

SAMPL Challenge

SAMPL Challenge is a chemistry topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand SAMPL Challenge rather than just read about it. In short: SAMPL (Statistical Assessment of the Modeling of Proteins and Ligands) is a set of community-wide blind challenges aimed to advance computational techniques as standard predictive tools in rational drug design. A broad range of biologically relevant systems with different sizes and levels of complexities including proteins, host–guest complexes, and drug-like small molecules have been selected to test the latest mod…

Key takeaways

  • SAMPL Challenge belongs to chemistry; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect SAMPL Challenge to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of SAMPL Challenge from memory before moving on to harder problems.

Reference excerpt

SAMPL (Statistical Assessment of the Modeling of Proteins and Ligands) is a set of community-wide blind challenges aimed to advance computational techniques as standard predictive tools in rational drug design. A broad range of biologically relevant systems with different sizes and levels of complexities including proteins, host–guest complexes, and drug-like small molecules have been selected to test the latest modeling methods and force fields in SAMPL. New experimental data, such as binding affinity and hydration free energy, are withheld from participants until the prediction submission deadline, so that the true predictive power of methods can be revealed. The most recent SAMPL5 challenge contains two prediction categories: the binding affinity of host–guest systems, and the distribution coefficients of drug-like molecules between water and cyclohexane. Since 2008, the SAMPL challenge series has attracte interest from scientists engaged in the field of computer-aided drug design (CADD) The current SAMPL organizers include John Chodera, Michael K. Gilson, David Mobley, and Michael Shirts.

Project significance The SAMPL challenge seeks to accelerate progress in developing quantitative, accurate drug discovery tools by providing prospective validation and rigorous comparisons for computational methodologies and force fields. Computer-aided drug design methods have been considerably improved over time, along with the rapid growth of high-performance computing capabilities. However, their applicability in the pharmaceutical industry are still highly limited, due to the insufficient accuracy. Lacking large-scale prospective validations, methods tend to suffer from over-fitting the pre-existing experimental data. To overcome this, SAMPL challenges have been organized as blind tests: each time new datasets are carefully designed and collected from academic or industrial research laboratories, and measurements are released shortly after the deadline of prediction submission. Researchers then can compare those high-quality, prospective experimental data with the submitted estimates. A key emphasis is on lessons learned, allowing participants in future challenges to benefit from modeling improvements made based on earlier challenges. SAMPL has historically focused on the properties of host–guest systems and drug-like small molecules. These simply model systems require considerably less computational resources to simulate than protein systems, and thus converge more quickly. Through careful design, these model systems can be used to focus on one particular or a subset of simulation challenges. The past several SAMPL host–guest, hydration free energy and log D challenges revealed the limitations in generalized force fields, facilitated the development of solvent models, and highlighted the importance of properly handling protonation states and salt effects.

Participation Registration and participation is free for SAMPL challenges. Beginning with SAMPL7, challenge participation data was posted on the SAMPL website, as well as the GitHub page for the specific challenge. Instructions, input files and results were then provided through GitHub (earlier challenges provided content primarily through D3R for SAMPL4-5, and via other means for earlier SAMPLs). Participants were allowed to submit multiple predictions through the D3R website, either anonymously or with research affiliation. Since the SAMPL2 challenge, all participants have been invited to attend the SAMPL workshops and submit manuscripts to describe their results. After a peer-review process, the resulting papers, along with the overview papers which summarize all submitting data, were published in the special issues of the Journal of Computer-Aided Molecular Design.

Funding The SAMPL project was recently funded by the NIH (grant GM124270-01A1), for the period of Sept. 2018 through August 2022, to allow the design of future SAMPL challenges to drive advances in the areas they are most needed for modeling efforts. The effort is spearheaded by David L. Mobley (UC Irvine) with co-investigators John D. Chodera (MSKCC), Bruce C. Gibb (Tulane), and Lyle Isaacs (Maryland). Currently challenges and workshops are run in partnership with the NIH-funded Drug Design Data Resource, but this will likely change over time as funding for the two projects is not coupled. Funding also allowed a broadening the scope of SAMPL; through SAMPL6, its role had been seen as primarily focused on physical properties, with D3R handling protein-ligand challenges. However, the funded effort broadened its focus to include systems which will drive improvements in modeling, including potentially suitable protein-ligand systems. This is still in contrast to D3R, which relies on donated datasets of pharmaceutical interest, whereas SAMPL challenges are specifically designed to focus on specific modeling challenges.

History

Earlier SAMPL challenges The first SAMPL exercise, SAMPL0 (2008) focused on the predictions of solvation free energies of 17 small molecules. A research group at Stanford University and scientists at OpenEye Scientific Software carried out the calculations. Despite the informal format, SAMPL0 laid the groundwork for the following SAMPL challenges. SAMPL1 (2009) and SAMPL2 challenges (2010) were organized by OpenEye and continued to focus on predicting solvation free energies of drug-like small molecules. Attempts were also made to predict binding affinities, binding poses and tautomer ratios. Both challenges attracted significant participations from computational scientists and researchers in academia and industry.

… excerpt ends here. Continue reading the full article.

Worked examples

Example 1 — a first encounter with SAMPL Challenge

Start with the simplest possible case. Write down what SAMPL Challenge claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In chemistry, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to SAMPL Challenge before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about SAMPL Challenge ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of SAMPL Challenge

In research
SAMPL Challenge appears in chemistry research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses SAMPL Challenge in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
SAMPL Challenge is common in secondary-school and first-year university syllabi. It links to neighbouring topics Computational chemistry, Drug discovery, so understanding it makes those chapters shorter.
In everyday life
Look for SAMPL Challenge outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study SAMPL Challenge in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what SAMPL Challenge means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain SAMPL Challenge out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is SAMPL Challenge in simple terms?

SAMPL (Statistical Assessment of the Modeling of Proteins and Ligands) is a set of community-wide blind challenges aimed to advance computational techniques as standard predictive tools in rational drug design. A broad range of biologically relevant systems with different sizes and levels of comple…

Why does SAMPL Challenge matter?

Because it connects several chemistry ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study SAMPL Challenge?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on SAMPL Challenge.

Tags

  • Computational chemistry
  • Drug discovery

Keep exploring