lunes, 28 de marzo de 2016

Optimization of continuous characters in parsimony




The node state in a phylogenetic tree can be easily calculate using discrete characters considering that the state is one option among a finite number of possibilities. However, calculate a node state using continuous characters is very complicated because the numbers of states are infinite (Goloboff et al, 2006). For this reason, continuous characters must be discretized. Farris (1970) proposes to treat continuous characters as additive characters. Several algorithms have been implemented (Goloboff, 1993 ; Thiele, 1993) to discretized continuous characters. Thus, I evaluated a Goloboff's algorithm and Thiele's algorithm by the Robinson-Foulds Metrics and Consistency Index (CI) expecting that the level of homoplasy decreases using continuous characters added to discrete characters. In order to, I simulated a tree of 25 tips (Figure 1) using geiger R library (Harmon et al, 2008), also simulated 3 different types of continuous characters and DNA sequence with 500 bp, model JC using seq-gen (Rambaut and Grassly, 2001). I recovery the tree in TNT (Goloboff, 2008) using only continuous characters, only discrete characters, continuous and discrete characters, different numbers (3, 10, 20) of states for thiele's algorithm and Goloboff algorithm, after I estimate the CI (Figure 2, Figure 3, Figure 4) and compare the trees with the original using Robinson-Foulds metric (Steel and Penny 1993) (Table 1).


Figure 1. Used simulated tree for generate continuous and discrete data.


Figure 2. Consistency Index for continuous (continuos) and discretes characters (discretos)

This could be because only three continuous characters were used. However, when more discrete continuous characters were used the degree of homoplasy decreased (Figure 3), this is consistent with the results obtained by Goloboff et al (2006).


 
Figure 3. Consistency Index for discrete characters (discretos) and continuous plus discrete characters (mix goloboff).


Figure 4. Consistency Index for total evidence using the Goloboff algorithm and thiele algorith with 3, 10 and 20 numbers of states.


Table 1. Robinson-Foulds Metric.





Goloboff algorithm yielded best results in the optimization way of characters. The problem with the Thiele’s algorithm is that always is subject to the subjectivity of the choice of the number of possible states. However, the Goloboff’s algorithm is not necessarily the best option for the reconstruction of phylogenies using total evidence. Robinson-Foulds metric (Table 1) showed that the smallest difference between the tree recovered and the original was obtained using the Thiele’s algorithm.

A copy of the scripts and used data can be found in https://github.com/dpabon/bio_comparada/tree/master/continuous_parsimony 

References

Farris, J., 1970. Methods for computing Wagner trees. Syst. Zool. 19,83–92. 
Goloboff, P., 1993b. Character optimization and calculation of tree lengths. Cladistics 9, 433–436.
Goloboff, Pablo A., Camilo I. Mattoni, and Andrés Sebastián Quinteros.' Continuous Characters Analyzed as Such’. Cladistics 22, no. 6 (December 2006): 589–601. doi:10.1111/j.1096-0031.2006.00122.x.
Goloboff, Pablo A., James S. Farris, and Kevin C. Nixon. ‘TNT, a Free Program for Phylogenetic Analysis’. Cladistics 24, no. 5 (1 October 2008): 774–86. doi:10.1111/j.1096-0031.2008.00217.x.

Harmon Luke J, Jason T Weir, Chad D Brock, Richard E Glor, and Wendell   Challenger. 2008. GEIGER: investigating evolutionary radiations. Bioinformatics 24:129-131.

Rambaut, A. and Grassly, N. C. (1997) Seq-Gen: An application for the Monte Carlo simulation of DNA sequence evolution along phylogenetic trees. Comput. Appl. Biosci. 13: 235-238.

Steel M. A. and Penny P. (1993) Distributions of tree comparison metrics - some new results, Syst. Biol.,42(2), 126-141

Thiele, K., 1993. The Holy Grail of the perfect character: the cladistic
treatment of morphometric data. Cladistics 9, 275–304.

viernes, 18 de marzo de 2016

Statistical consistency of Maximum Likelihood and Maximum Parsimony

Statistical consistency is often one of the factors often cited to favor statistical estimation methods in phylogenetic reconstruction . It is defined as the porperty of estimator, for wich as data (characters)  tends to infinity, the inferred topology lead to  converge to the true: the real phylogeny (Truszkowski, & Goldman, 2015).

Debate about inconsistency  has a great story behind it. Since the proposal Felsenstein(1978) with a simple evolutionary model, much work has been done to evaluate the statistical properties of estimators in phylogeny under a set of different scenarios. Much of the effort has made to show that model based approaches like Maximum Likelihood (and Bayesian Inference) has advantages over Parsimony (Wheeler, 2011).

Here I test the statistical consistency of Maximum Likelihood  and two flavours of parsimony reconstruction in different sets of  topologies and models of nucleotide substitution,  with four taxa-trees  in a simulated experiment.

Materials and methods

To carry out the analysis of statistical consistency. First nucleotide sequences were simulated in the software Seq-Gen (Rambaut & Grass, 1997). in three different models of DNA evolution: GTR, and JC HKY. Also  each model was analyzed in four topologies (original topologies) whose length (q and p) differs as defined below:


  • Farris Zone.  q=0.6, p=0.1.Tips with long branches are sisters .
  • Felsenstein Zone. p=0.6, q= 0.1.  Tips with branches of different length are sisters.
  • Long Branches. p=q =0.6
  • Short Branches.p=q= 0.1
  • Central branch of topologies always take the value of q.

For maximum parsimony analysis two approaches were used: Static and dynamic homology (The static homology parsimony was conducted in TNT (Goloboff et. al, 2008). POY ( Varón et al, 2010) was the software used in the analysis of dynamic homology. Maximum Likelihood inferences were carried out in PhyML software (Guindon et al, 2010).


To asses differences between the inferred topologies and the originals, the symmetric difference (Steel & Penny, 1993) was used as implemented in RF.dist function ot the R package  phangorn (Schliep, 2011). The frequency of the true tree recovered  was the measure of consistency here used.

Results and Discussion.

Results show that  maximum parsimony analysis are consistent in three of the four areas evaluated or topologies. These are the Farris zone, short branches and  Long Branches. Maximum likelihood was equally consistent in three zones: Long branches, short branches and Felsenstein zone (Figs 1,2,3).
There is no dramatic difference between frequencies of recovered rigth trees by model.  

The tendency is to converge to the rigth tree in three of the four zones as the data avalaible increase, regardless the model of the data.



 Fig 1.  Frequency of recovered trees under the JC model. X axis show log of  base pairs. 
Fig.2.  Frequency of recovered trees under the HKY model. X axis show log of  base pairs.



Fig 3. Frequency of recovered trees under the GTR model. X axis show log of  base pairs. 

It can be seen as direct optimization leads to inconsistency much faster than does the traditional static parsimony under the JC model. These results confirm that, regardless of the concept of homology used,  parsimony  is inconsistent (Warnow, 2012) for the  Jukes Cantor  model  and more complex models in the Felsenstein zone.
 Statistical consistency is a desirable property in estimates of phylogenetic reconstruction (Wheeler, 2011). Under different configurations branches estimators behave similarly and all fail in at least some zones, even under simple models.
If the three (or two) approaches here assesed show a similar behavior (Failing to achieve perfect consistency in at least one zone), and perfect consistency is not reached one must choose an approach between them with freedom knowing that in some configuration of branch length it can be yield the rigth tree when the avalaible data is enougth.


References

Felsenstein, J. (1978). Cases in which parsimony or compatibility methods will be positively misleading. Systematic Biology, 27(4), 401–410.

Goloboff, P. A., Farris, J. S., & Nixon, K. C. (2008). TNT, a free program for phylogenetic analysis.Cladistics, 24(5), 774-786.


Guindon, S., Dufayard, J. F., Lefort, V., Anisimova, M., Hordijk, W., & Gascuel, O. (2010). New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Systematic biology, 59(3), 307-321


Rambaut, A., & Grass, N. C. (1997). Seq-Gen: an application for the Monte Carlo simulation of DNA sequence evolution along phylogenetic trees.Computer applications in the biosciences: CABIOS, 13(3), 235-238.

Schliep, K. P. (2011). phangorn: Phylogenetic analysis in R. Bioinformatics, 27(4), 592–593.

Steel, M. A., Hendy, M. D., & Penny, D. (1993). Parsimony can be consistent! Systematic Biology, 581–587.

Truszkowski, J., & Goldman, N. (2015). Maximum likelihood phylogenetic inference is consistent on multiple sequence alignments, with or without gaps.Systematic biology, syv089.

Varón, A., L. S. Vinh, W. C. Wheeler. 2010. POY version 4: phylogenetic analysis using dynamic homologies. Cladistics, 26:72-85.

Warnow T. Standard maximum likelihood analyses of alignments with gaps can be statistically inconsistent. PLOS Currents Tree of Life. 2012 Mar 12 . Edition 1. doi: 10.1371/currents.RRN1308.

Wheeler, W. C. Comparison of Optimality Criteria. Systematics: A Course of Lectures, 2011.269-287.

ML and IB consistency

Statistical consistency is the capacity of the estimator that be true when the lenght of data tend to infinite. Felsenstein (1978) show that phylogenetic reconstruction using parsimony is incosistent when the branch lengths have certain proportions (Figure 1) this proportions have the name of Felsentein zone. However Farris (1999) show that phylogenetic reconstruction using Maximum Likelihood (ML) is inconsistent also when branch lengths have certain proportions (Figure 2) this type of zone have the name of Farris zone.
Contrary to the previous phylogenetics reconstructions techniques, bayesian analysis (IB) is consistency without import the zone or branch lengths relations (Steel, 2013).

Figure 1. Felsenstein zone


Figure 2. Farris zone


However i checked experimentally if is true the consistency of IB and ML. For this i use four type of unrooted trees with 4 tips (farris zone, felsenstein zone, "long branches", "small branches"). I choose three sizes base pairs (100, 1000, 10000) and for each size i generated 25 sequences of DNA based in the tree and model JC . I recover the trees using Mrbayes v 3.2.6 compiled to MPI and Phyml 20160303 also compiled to MPI, and i using the JC model. I compared each tree vs original tree using Robinson Fould metric (Steel & Penny, 1993). I define the consistency as: Number of trees that Robinson Fould metric = 0/Number of total trees for each type of tree and size of bases pairs.  
In the Farris zone (Figure 3) i need sample more sizes of DNA because only three sizes show that ML is consistent in the Farris zone and this is false, IB was consistent in this zone. In the Felsentein zone (Figure 4), ML and IB were consistent. In "long branches"(Figure 5) and "small branches" (Figure 6) ML and IB were consistent. For this is true that IB is consistency independently of the zone or branches lenght proportions. 






Figure 3. Statistical consistency in Farris zone. ML = Maximum Likelihood, IB =Bayesian Inference

Figure 4. Statistical consistency in Felsenstein zone. ML = Maximum Likelihood, IB =Bayesian Inference

 
Figure 5. Statistical consistency for "Long branches". ML = Maximum Likelihood, IB =Bayesian Inference

 
I programmed some useful functions in R for the threes.
 
https://github.com/dpabon/bio_comparada/blob/master/consistencia/bin/R/functions.R
 https://github.com/dpabon/bio_comparada/blob/master/consistencia/bin/R/gentree.R
If you want run all analysis in your computer you can download a copy of the repository and follow the instructions.

https://github.com/dpabon/bio_comparada/tree/master/consistencia

Farris, J. (1999). Likelihood and Inconsistency. Cladistics 15, 199–204

Felsenstein, J. (1978). Cases in which parsimony or compatibility methods will be positively misleading. Systematic Biology, 27(4), 401–410.

Steel M. A. and Penny P. (1993) Distributions of tree comparison metrics - some new results, Syst. Biol.,42(2), 126-141

Steel, M. (2013). Consistency of Bayesian inference of resolved phylogenetic trees. Journal of theoretical biology 336: 246–49.

domingo, 29 de noviembre de 2015

Phylogenetic inference and philosophy, different approaches for the same purpose


************** This post was updated on 27.01.2016***************


Phylogenetic inference attempts to elucidate the evolutionary relationships among organisms.
Various approaches have been made for this purpose, they differ in their rationale for addressing the problem (De Queiroz & Poe, 2001). Among the best known approaches are parsimony, bayesian inference and  likelihoodism. Below I will discuss some of its characteristics, basic assumptions and finally express which one, in my opinion, comes more adequately to face a phylogenetic analysis.

Parsimony has its grounds in the principle of simplicity. Proponents of the principle of parsimony argue that this approach is justified by the ideas of Karl Popper. This implies that the phylogenetic hypothesis must be falsifiable and  rigorous tests, so be corroborated. From this perspective, those hypotheses with the least amount of  changes should be preferred, thus minimizing the number of ad hoc explanations (Grupe & Harbeck, 2015). But the preference for simpler explanations does not mean that nature behaves well, evolution does not have to be parsimonious. This seems difficult to understand, and in fact, sounds contradictory.

Statistical approaches such as maximum likelihood or Bayesian inference share using evolutionary models that take into account the probability of  changes  between  character states and base frequencies (Archibald, Mort, & Crawford, 2003). In a parsimonious approach  seems that these changes are equally likely.

The likelihood method can be defined as the probability of a hypothesis given the data   of the data given a hypothesis. This approach seeks to find that hypothesis explains the observed data (characters) in the manner that maximizes the probability that these are observed.

Bayesian inference is different from the likelihood that takes into account prior knowledge to the observations,  it is posibe assign probabilities to hypotheses (Topologies) before observations are made. The main problem with the Bayesian inference is its distinguishing feature. The priors can be a double-edged sword, on the one hand allow the process to incorporate prior knowledge of phylogenetic inference, but actually priors are difficult to accurately estimate  (Velasco,2008). For several authors is a common practice then assign equal priors, but this means that the main advantage of this method had just wasted . Additionally, what if the priors are estimated incorrectly, this could result in a bias in the results of the process.

The advantages of statistical methods seem obvious, they make assumptions on a given model. Parsimony however, assumes any evolutionary model, or does this mean that the changes are equally likely, it does not seem logical, especially if we speak of continuous characters. Given these difficulties with the priors, in my opinion, a likelihoodism approach is most appropriate for phylogenetic inference.

References


  • Grupe, G., & Harbeck, M. (2015). Taphonomic and Diagenetic Processes. En W. Henke & I. Tattersall (Eds.), Handbook of Paleoanthropology (pp. 417–439).
  •  De Queiroz, K., & Poe, S. (2001). Philosophy and phylogenetic inference: a comparison of likelihood and parsimony methods in the context of Karl Popper’s writings on corroboration. Systematic Biology, 50(3), 305–321.
  • Archibald, J. K., Mort, M. E., & Crawford, D. J. (2003). Bayesian inference of phylogeny: a non-technical primer. Taxon, 187–191.
  • Velasco, J. D. (2008). Philosophy and The Tree of Life (Doctoral dissertation, Ph. D. Thesis). University of Wisconsin-Madison).

  




Maximum Likelihood versus Parsimony and Bayesian inference

Many authors emphasize in parsimony method by resorting to realism, and the simplicity of the assumptions (Goloboff, 2003). While others say that parsimony subject to specific models is the same as Likelihood (Farris, 1983). Below I will discuss some arguments that lower use of parsimony as a method to clarify the evolutionary relationships among organisms. 

First, that offers simplicity parsimony in assumptions does not mean that these are clear and they are the best. Many wonder what really are the assumptions about the evolutionary process that takes parsimony method? Just assume that the offspring having modification with respect to their ancestors? Assumptions parsimony leave many doubts. 

Second, parsimony does not discriminate changes in the branches are more probability or improbable. It not assumed if a branch is more probability to change over another (Sober, 2004). It is, for parsimony no selection for one character over another (Sober, 2002). It is at this point that the Maximum Likelihood method has its advantages. Using evolutionary models allows us from propositions given by the data and calculate the probability given the hypothesis (Goloboff, 2003). In addition to the rate of Likelihood we can measure the strength of the statistical evidence and so choose the topology more Likelihood (Royall, 1999). 

 On the other hand, it is the Bayesian inference method, a probabilistic method like Likelihood uses evolutionary models. This method uses priors basis for calculating the posterior probability of the data given hypothesis. One risk of using priors is that these can become subjective and condition the calculation of posterior probabilities. I personally think that Bayesian inference is a modification of Likelihood, but with more potential for bias given the priors. 


Andrea Lizeth Silva Cala 

Reference

Goloboff, P. A. (2003). Parsimony, likelihood, and simplicity. Cladistics, 19(2), 91-103.

Farris, J. (1983). The logical basis of phylogenetic analysis (pp. 7-36). na.

Sober, E. (2004). The contest between parsimony and likelihood. Systematic biology, 53(4), 644-653.

Sober, E. (2002): “Reconstructing Ancestral Character States – A Likelihood Perspective on Cladistic Parsimony.” The Monist 85: 156-176.

Royall, R. (1999): The Strength of Statistical Evidente. Johns Hopkins University Department of Biostatisks 615 North Wolfe Street Baltimore MD 2120.5 USA

Bayesianism and likelihoodism


¿Bayesianism or Likelihoodism?

Let me start with the Royall's three questions:

1. ¿What does the present evidence say?
2. ¿What should you believe?
3. ¿What should you do?

Although Likelihoodists and Bayesians both share the likelihood principle and the law of likelihood which are important in the philosophy of scientific method, they disagree on several instances:

Its necessary highlight that the most remarkable difference between them is that Bayesians use prior probabilities in other words posterior probability distributions that require prior probability distributions and likelihood functions and likelihoodists not.

Another difference points to the meaning of evidence: Likelihoodists characterize data as evidence and they don't use them to guide our beliefs or actions and maintain that this characterization is valuable in itself (Royall 1997, Ch. 1). Then, you couldn't give answer to the second nor the third question of Royall because they say nothing about what you should believe after you receiving the evidence without take into account what you believe before receiving the evidence. On the other hand forBayesians the prior is updated in the light of new data that is the evidence (Sober, 2008) from this perspective you could give answer to all questions.

Regarding to the second question about your degree of belief Bayesians answer this question from the concept of confirmation where the observation (O) provides confirmation of hypothesis 1 (H1) when this has a higher likelihood than its own negation (Gandenberger, 2013). Unlike Likelihoodists whom doesn't use this concept of confirmation, they don't take into account if the evidence raises, lowers or not change the probability of the hypothesis. They compare hypotheses to each other which have their own likelihoods and use the law of likelihood to interpret the data where: the observation (O) favors hypothesis 1 (H1) over hypothesis 2 (H2) if Pr (O | H1) > Pr (O | H2) and the likelihood ratio is used to show the degree to which O favors H1 over H2 that is given by Pr(O | H1) / Pr(O | H2), and they ask if H1 has a higher likelihood than H2. So, for likelihoodists is enough use the likelihood ratio as a measure of degree favoring one hypothesis over other one (Sober, 2008). In contrast to Bayesian for whom is not enough and then implemented the use of posterior probabilities, see below.

In Bayesian inference you assign a probability to the hypothesis (H) before doing an observation in other words is the distribution of the parameters before doing analysis of the data (prior probability) and after of doing it there is a reallocation of the probability assigned to H and the probability in the light of evidence is known as posterior probability and is denoted Pr(H|O) that means probability of the Hypothesis given the Observation. In contrast Maximum likelihood where the likelihood of the hypothesis is the probability that H confers on O Pr(O|H) (Sober, 2008).

Other thing in common is that ML and BI use the same models of evolution, but the way to measure the support of relationships in the topology are different, ML uses bootstrap support (BS) which is a measure of confidence, and uses data resampling to estimate the support (Cummings et al., 2003). Unlike BI that uses the posterior probability (PP) which is calculated from prior probability, likelihood functions and data. Both measures have been controversial because of several reasons and some claim there is a equivalence between both measures (Efron, H. and Holmes, 1996), but some studies like Erixon et al. (2003) reject this assumption and others claim PP is a better measure of support (Alfaro, Zoller, and Lutzoni, 2003).

Given the similarities and differences between them I think that Bayesian inference is the best method of all.

Bibliography

Alfaro, M. E., Zoller, S., & Lutzoni, F. (2003). Bayes or bootstrap? A simulation study comparing the performance of Bayesian Markov chain Monte Carlo sampling and bootstrapping in assessing phylogenetic confidence. Molecular Biology and Evolution, 20(2), 255–266.

Cummings, M. P., Handley, S. A., Myers, D. S., Reed, D. L., Rokas, A., & Winka, K. (2003). Comparing bootstrap and posterior probability values in the four-taxon case. Systematic Biology, 52(4), 477–487.

Efron, B., Halloran, E., & Holmes, S. (1996). Bootstrap confidence levels for phylogenetic trees. Proceedings of the National Academy of Sciences, 93(23), 13429.

Erixon, P., Svennblad, B., Britton, T., & Oxelman, B. (2003). Reliability of Bayesian posterior probabilities and bootstrap frequencies in phylogenetics. Systematic Biology, 52(5), 665–673.

Gandenberger Greg . 2013. Why I am not a likelihoodist.


Royall, R. Statistical Evidence: A Likelihood Paradigm, Boca Raton, Fla.:Chapman and Hall.(1997).

SOBER, Elliott. Evidence and evolution: The logic behind the science. Cambridge University Press, 2008.


 

Philosophy in the biological world

The discussion about construction of how knowledge is built are not new, from Plato and his proposed world of ideas to Kant in his criticism to pure reasons(1) is underlined that the construction of knowledge which we call science is not dogmatic and static, instead it is a non volatile element and  a conditioned subject to time-space paradigm own of humanity, reducing everything to a purely linguistic problem (2). Therefore, syncretism is not a symptom of intellectual immaturity or inferiority of it, but a prudent demonstration against scholastic thought own religious processes, some scholars confused with scientific work.

Added to all this it is important to clarify that, contrary to sciences like mathematics, physics, and chemistry. The theoretical corpus in biological sciences completely lacks axiomatic systems that support the developed theories. While evolution is a fact. The causality of the phenomenon is highly debated due to the number of ad-hoc theories and hypotheses (3, 4). It is in this environment that the phylogenetic theories that attempt to answer the evolutionary relationships of organisms are developed. Therefore this essay  put on the table Bayesian analysis , the likelihood and parsimony as "irreconcilable" philosophical. Trying to approach the reality of evolutionary phenomena (not yet finished to be clear) to consider some as "the best ".

Let's start with the parsimony in which methodologically the tree with fewer transformations is selected, this is derivative of the philosophical principle that the simplest theory must be correct for being the least complex (5). A strongly nominalism position that can result in multiple ad- hoc theories that could never be applied to different events rather than the themselves cases. Following this logic we could generate own theories for each type of phylogenetic relationships of every living form. Which will lead to a greater number of hypotheses the number of species (counting the species already extinct). Which paradoxically ends up being contrary to the principle of parsimony.
In contrast to the  parsimony, the probabilistic models certainly have an advantage developing stochasticity in their methods, thus avoiding a possible fall in phylogenetic Laplace demon advantage. Maximum likelihood analyzes the conditional probability of the observations given the hypothesis, or in more colloquial terms how well the data fit in a given hypothesis. In which each tree is considered a hypothesis generated by choosing the highest likelihood. However likelihood ignores previously gathered evidence about the event at that epistemological terms can only be assigned certain degree of value as truth. By contrast,  Bayesian analysis is an excellent tool for the analysis of phylogenetic hypothesis by quantifying past evidence (prior) and included in the phylogenetic analysis. It is therefore the most appropriate tool to address the problem to elucidate the evolutionary relationships of living forms.

References

  1. Kant, Immanuel, and Norman Kemp Smith. 1929. Immanuel Kant's Critique of pure reason. Boston: Bedford.
  2. Wittgenstein, Ludwig. 1922. Tractatus logico-philosophicus. London: Routledge & Kegan Paul.
  3. Margulis, Lynn; Dorion Sagan (2003). Captando Genomas. Una teoría sobre el origen de las especies. Ernst Mayr (prólogo). David Sempau (trad.) (1ª edición). Barcelona: Editorial Kairós
  4. Darwin, Charles Robert. The Origin of Species. Vol. XI. The Harvard Classics. New York: P.F. Collier & Son, 1909–14.
  5. Robert Audi, ed., Ockham's razor, The Cambridge Dictionary of Philosophy (2nd Edition), Cambridge University Press.