ArticleslgStudy

biology

PubChem

PubChem is a biology topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand PubChem rather than just read about it. In short: PubChem is a database of chemical molecules and their activities against biological assays. The system is maintained by the National Center for Biotechnology Information (NCBI), a component of the National Library of Medicine, which is part of the United States National Institutes of Health (NIH).

Key takeaways

  • PubChem belongs to biology; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect PubChem to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of PubChem from memory before moving on to harder problems.

Reference excerpt

PubChem is a database of chemical molecules and their activities against biological assays. The system is maintained by the National Center for Biotechnology Information (NCBI), a component of the National Library of Medicine, which is part of the United States National Institutes of Health (NIH). PubChem can be accessed for free through a web user interface. Millions of compound structures and descriptive datasets can be freely downloaded via FTP. PubChem contains multiple substance descriptions and small molecules with fewer than 100 atoms and 1,000 bonds. More than 80 database vendors contribute to the growing PubChem database.

History PubChem was released in 2004 as a component of the Molecular Libraries Program (MLP) of the NIH. As of November 2015, PubChem contains more than 150 million depositor-provided substance descriptions, 60 million unique chemical structures, and 225 million biological activity test results (from over 1 million assay experiments performed on more than 2 million small-molecules covering almost 10,000 unique protein target sequences that correspond to more than 5,000 genes). It also contains RNA interference (RNAi) screening assays that target over 15,000 genes. As of August 2018, PubChem contains 247.3 million substance descriptions, 96.5 million unique chemical structures, contributed by 629 data sources from 40 countries. It also contains 237 million bioactivity test results from 1.25 million biological assays, covering >10,000 target protein sequences. As of 2020, with data integration from over 100 new sources, PubChem contains more than 293 million depositor-provided substance descriptions, 111 million unique chemical structures, and 271 million bioactivity data points from 1.2 million biological assays experiments.

Databases PubChem consists of three dynamically growing primary databases. As of 5 November 2020 (number of BioAssays is unchanged):

Compounds, 111 million entries (up from 94 million entries in 2017), contains pure and characterized chemical compounds. Substances, 293 million entries (up from 236 million entries in 2017 and 163 million in Sept. 2014), contains also mixtures, extracts, complexes and uncharacterized substances. BioAssay, bioactivity results from 1.25 million (up from 6,000 in Sept. 2014) high-throughput screening programs with several million values.

Searching Searching the databases is possible for a broad range of properties including chemical structure, name fragments, chemical formula, molecular weight, XLogP, and hydrogen bond donor and acceptor count. PubChem contains its own online molecule editor with SMILES/SMARTS and InChI support that allows the import and export of all common chemical file formats to search for structures and fragments. Each hit provides information about synonyms, chemical properties, chemical structure including SMILES and InChI strings, bioactivity, and links to structurally related compounds and other NCBI databases like PubMed. In the text search form the database fields can be searched by adding the field name in square brackets to the search term. A numeric range is represented by two numbers separated by a colon. The search terms and field names are case-insensitive. Parentheses and the logical operators AND, OR, and NOT can be used. AND is assumed if no operator is used. Example (Lipinski's Rule of Five):

0:500[mw] 0:5[hbdc] 0:10[hbac] -5:5[logp]

Database fields

See also Chemical database CAS Common Chemistry - run by the American Chemical Society Comparative Toxicogenomics Database - run by North Carolina State University ChEMBL - run by European Bioinformatics Institute ChemSpider - run by UK's Royal Society of Chemistry DrugBank - run by the University of Alberta IUPAC - run by Swiss-based International Union of Pure and Applied Chemistry (IUPAC) Moltable - run by India's National Chemical Laboratory PubChem - run by the National Institute of Health, USA BindingDB - run by the University of California, San Diego SCRIPDB - run by the University of Toronto, Canada National Center for Biotechnology Information (NCBI) - run by the National Institute of Health, USA Entrez - run by the National Institute of Health, USA GenBank - run by the National Institute of Health, USA

References

External links

Official website

Worked examples

Example 1 — a first encounter with PubChem

Start with the simplest possible case. Write down what PubChem claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In biology, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to PubChem before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about PubChem ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of PubChem

In research
PubChem appears in biology research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses PubChem in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
PubChem is common in secondary-school and first-year university syllabi. It links to neighbouring topics Biological databases, Chemical databases, Public-domain software with source code, so understanding it makes those chapters shorter.
In everyday life
Look for PubChem outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study PubChem in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what PubChem means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain PubChem out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is PubChem in simple terms?

PubChem is a database of chemical molecules and their activities against biological assays. The system is maintained by the National Center for Biotechnology Information (NCBI), a component of the National Library of Medicine, which is part of the United States National Institutes of Health (NIH).

Why does PubChem matter?

Because it connects several biology ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study PubChem?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on PubChem.

Tags

  • Biological databases
  • Chemical databases
  • Public-domain software with source code
  • United States National Library of Medicine

Keep exploring