ArticleslgStudy

computer science

Machine learning in bioinformatics

Machine learning in bioinformatics is a computer science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Machine learning in bioinformatics rather than just read about it. In short: Machine learning in bioinformatics is the application of machine learning algorithms to bioinformatics, including genomics, proteomics, microarrays, systems biology, evolution, and text mining. Prior to the emergence of machine learning, bioinformatics algorithms had to be programmed by hand; for problems such as protein structure prediction, this proved difficult.

Machine learning in bioinformatics — main illustration
Machine learning in bioinformatics — illustration

Key takeaways

  • Machine learning in bioinformatics belongs to computer science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Machine learning in bioinformatics to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Machine learning in bioinformatics from memory before moving on to harder problems.

Reference excerpt

Machine learning in bioinformatics is the application of machine learning algorithms to bioinformatics, including genomics, proteomics, microarrays, systems biology, evolution, and text mining. Prior to the emergence of machine learning, bioinformatics algorithms had to be programmed by hand; for problems such as protein structure prediction, this proved difficult. Machine learning techniques such as deep learning can learn features of data sets rather than requiring the programmer to define them individually. The algorithm can further learn how to combine low-level features into more abstract features, and so on. This multi-layered approach allows such systems to make sophisticated predictions when appropriately trained. These methods contrast with other computational biology approaches which, while exploiting existing datasets, do not allow the data to be interpreted and analyzed in unanticipated ways.

Tasks Machine learning algorithms in bioinformatics can be used for prediction, classification, and feature detection and selection. Classification or recognition algorithms output a categorical class, while prediction algorithms output a numerical valued feature. Algorithms differ in the type of process used to build the predictive models from data using analogies, rules, neural networks, probabilities, and/or statistics. Due to the exponential growth of information technologies and applicable models, including artificial intelligence and data mining, in addition to the access to ever-more comprehensive and open data sets, new and better information analysis techniques have been created, based on their ability to learn.

Approaches

Artificial neural networks Artificial neural networks in bioinformatics have been used for:

Comparing and aligning RNA, protein, and DNA sequences. Identification of promoters and finding genes from sequences related to DNA. Interpreting the expression-gene and micro-array data. Identifying the network (regulatory) of genes. Learning evolutionary relationships by constructing phylogenetic trees. Classifying and predicting protein structure. Molecular design and docking

Feature engineering The way that features, often vectors in a many-dimensional space, are extracted from the domain data is an important component of learning systems. In genomics, a typical representation of a sequence is a vector of k-mers frequencies, which is a vector of dimension 4 k {\displaystyle 4^{k}} whose entries count the appearance of each subsequence of length k {\displaystyle k} in a given sequence. Since for a value as small as k = 12 {\displaystyle k=12} the dimensionality of these vectors is huge (e.g. in this case the dimension is 4 12 ≈ 16 × 10 6 {\displaystyle 4^{12}\approx 16\times 10^{6}} ), techniques such as principal component analysis are used to project the data to a lower dimensional space, thus selecting a smaller set of features from the sequences.

Classification In this type of machine learning task, the output is a discrete variable. One example of this type of task in bioinformatics is labeling new genomic data (such as genomes of unculturable bacteria) based on a model of already labeled data.

Hidden Markov models Hidden Markov models (HMMs) are a class of statistical models for sequential data (often related to systems evolving over time). An HMM is composed of two mathematical objects: an observed state‐dependent process X 1 , X 2 , … , X M {\displaystyle X_{1},X_{2},\ldots ,X_{M}} , and an unobserved (hidden) state process S 1 , S 2 , … , S T {\displaystyle S_{1},S_{2},\ldots ,S_{T}} . In an HMM, the state process is not directly observed – it is a 'hidden' (or 'latent') variable – but observations are made of a state‐dependent process (or observation process) that is driven by the underlying state process (and which can thus be regarded as a noisy measurement of the system states of interest). HMMs can be formulated in continuous time. HMMs can be used to profile and convert a multiple sequence alignment into a position-specific scoring system suitable for searching databases for homologous sequences remotely. Additionally, ecological phenomena can be described by HMMs.

… excerpt ends here. Continue reading the full article.

Illustrations

Machine learning in bioinformatics: Visual representation of the self-attention mechanism within a single transformer attention head, showing the transformation into query, key, and value vectors and the weighted scoring process
Visual representation of the self-attention mechanism within a single transformer attention head, showing the transformation into query, key, and value vectors and the weighted scoring process
Machine learning in bioinformatics: Some bioinformatic applications[which?] of random forest
Some bioinformatic applications[which?] of random forest
Machine learning in bioinformatics: The growth of GenBank, a genomic sequence database provided by the National Center for Biotechnology Information (NCBI)
The growth of GenBank, a genomic sequence database provided by the National Center for Biotechnology Information (NCBI)
Machine learning in bioinformatics: A protein's amino acid sequence annotated with the protein secondary structure. Each amino acid is labeled as either an alpha helix, a beta-sheet, or a coil.
A protein's amino acid sequence annotated with the protein secondary structure. Each amino acid is labeled as either an alpha helix, a beta-sheet, or a coil.
Machine learning in bioinformatics: A DNA-microarray analysis of Burkitt's lymphoma and diffuse large B-cell lymphoma (DLBCL), which differences in gene expression patterns
A DNA-microarray analysis of Burkitt's lymphoma and diffuse large B-cell lymphoma (DLBCL), which differences in gene expression patterns

Worked examples

Example 1 — a first encounter with Machine learning in bioinformatics

Start with the simplest possible case. Write down what Machine learning in bioinformatics claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In computer science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Machine learning in bioinformatics before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Machine learning in bioinformatics ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Machine learning in bioinformatics

In research
Machine learning in bioinformatics appears in computer science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Machine learning in bioinformatics in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Machine learning in bioinformatics is common in secondary-school and first-year university syllabi. It links to neighbouring topics Bioinformatics, Machine learning, so understanding it makes those chapters shorter.
In everyday life
Look for Machine learning in bioinformatics outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Machine learning in bioinformatics in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Machine learning in bioinformatics means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Machine learning in bioinformatics out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Machine learning in bioinformatics in simple terms?

Machine learning in bioinformatics is the application of machine learning algorithms to bioinformatics, including genomics, proteomics, microarrays, systems biology, evolution, and text mining. Prior to the emergence of machine learning, bioinformatics algorithms had to be programmed by hand; for p…

Why does Machine learning in bioinformatics matter?

Because it connects several computer science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Machine learning in bioinformatics?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Machine learning in bioinformatics.

Tags

  • Bioinformatics
  • Machine learning

Keep exploring