lunes, 3 de marzo de 2008

On Bayesian Inference (BI)

Huelsenbeck etal. (2001) said Bayesian inference of phylogeny is a powerful tool for addressing a number of long-standing, complex questions in evolutionary biology. The power they talk about states in the plausibility of fixing prior distributions of the parameters, and based on those prior probabilities and the likelihood of the data inferred the posterior probabilities of a tree. The posterior probability of each clade is estimated based on the frequency at which that clade is recovered among sampled trees once stationary log-likelihood have been reached under an MCMC algorithm. The numbers on the branches are said to represent the probability that the clade is correct or true (Huelsenbeck etal. 2002). The MCMC chain also sums another virtue to the Bayesian inference: makes it fast.

But are real those virtues or is only a hope that attracts pushovers? Firstly the prior problem, how do we calculate the probability of something we have not seen (Sober, 2002)? Someone could say that the answer to the trees prior distribution is not a problem and we can use the same probability to all of them (flat priors) and consider all possibilities, but this takes off one of the magnificent virtues of the Bayesian inference that is to involve into the calculation the plausibility of an event to occur. Although, Steel & Pickett (2006) proved that only under a special case priors do not induces a uniform distribution on clades which makes impossible the support evaluation for particular clades when the probability could be influenced by the clade size.

One of the most attractive features of the Bayesian inference is the speed, but when we need to be sure that the chains converge the time increment with the complexity and size of the data sets (Goloboff and Pol., 2005). On chains convergence roots the possibility of estimate the posterior probability of each clade, so a wrong implementation of the method drives to mistaken estimations. Moreover, admitting that the posterior probability has being well estimated the way the MCMC chain is summarized could give auto-inconsistent answers (the topology with the data) because the majority rule consensus may not recognize certain similarities among trees, and may be a poor summary (Yang, 2006). Besides, the posterior probability cannot be seen as a universal probability of truth, because it is given by the data, the model and the prior, so it is just a “local” probability (Simmons etal., 2004; Yang, 2006). Finally, but not less, actually Bayesian inference inflates the probabilities of correct clades and recovers high probabilities for incorrect nodes ( i.e. Douody etal., 2003; Simons etal., 2004; Goloboff & Pol., 2005).

On IB

Bayesian methods deal with the notion of a probability distribution for the parameter; the distribution of the parameter before the data are analyzed is called the prior distribution. The Bayes theorem is used in Bayesian inference (BI) to calculate the posterior distribution of the parameter, that is, the conditional distribution of the parameter given the data (Holder and Lewis, 2003).

The general idea beyond the “tree inference” done by IB is to construct a Markov Chain that has as its state space the parameters of the statistical model, and a stationary distribution that is the posterior probability distribution of the parameters and run a sampling chain for long enough time; then sort sampled trees in probability order and pick trees until cumulative probability is reached! (Yang, 2006). Clearly it is not conceived as a search mechanism, but instead as a sampling mechanism and therefore it probably won’t find the individual trees of maximum a posteriori probability. There is still the difficulty of when to know whether the chain has run long enough and when the method converges, facts that are ignore in many publications.

There is an idea that IB provides measures of support faster than ML bootstrapping. Bayesian inference produces both a tree estimate and measures of uncertainty for the groups on the tree (Holder and Lewis, 2003). But the problem is that it attribute a high probability to false groups that should at least be recognized as ambiguous (Albert, 2005) and, indeed, when recognizing monophyletic group IB does it a frequentist way: check how many sampled trees claim a particular group is monophyletic and this is the probability of our group of being monophyletic!.

The optimal hypothesis under BI is the one that maximizes the posterior probability (Holder and Lewis, 2003). The posterior probability for a hypothesis is proportional to the likelihood multiplied by the prior probability of that hypothesis. In many publications Prior probabilities are ignored or a uniform distribution over the range of the parameters (flat priors) is used. It have been suggested that priors can be specified by using either an objective assessment of prior evidence concerning the parameter or the researcher's subjective opinion! (Yang, 2006), but when no available information about the parameter exists, it is unclear which prior is more reasonable. In the other hand it has been shown that uniform priors are not non informative -no prior represent total ignorance- and is generally accepted that personal prejudices influence statistical inference.


Note on Support:

If a data set contains homoplasy then different characters support different trees, hence which tree (or trees) a given data supports will depend on which characters have been sampled (Page and Holmes, 1998). I consider that support is a measure of how perturbation in the data gives a different result given that repeated sampling from the population is difficult and sometimes we are interested in what we call repeatability: the probability that another such sample shares the groups with the original one.

Estimates of phylogeny based on samples will be accompanied by sampling error. One way to measure sampling error is to take multiple resamples (pseudoreplicates) from our sample and build a tree. The variation among estimates derived from each pseudoreplicate is a measure of the sampling error associated with our sample. The simple bootstrap can be applied therefore as a perturbation tool to asses the stability (in the sense of continuity, a small perturbation in the data that produces only a small perturbation in the data that produces only small perturbation in the estimate) of the estimator (Holmes, 2003). Bootstrapping and jackknife (Bayesian methods based on Markov chain Monte Carlo as well) essentially make confidence statements for the trees. The other approach, Bremer support, examines how many extra steps are needed to lose a branch in the consensus tree of near-most-parsimonious trees. This method explores suboptimal solutions and determines how much worse a solution must be for a hypothesized group not to be recovered -the amount of contradictory evidence required to refute a group (Bremer, 1994)-.

On Bayesian Inference and Support

Bayesian Inference


The Bayesian inference (BI) in Systematics (Rannala & Yang, 1996; Yang & Rannala, 1997) is a most frecuently used methods in phylogenetic analysis in the last decade. The speed of BI in comparision with other methods such as Parsimony and Maximum Likelihood (ML). Further, the posteriori probability is a attractive concept to show certainty of the results in an analysis. However, BI posses several problems and mistakes in the phylogenetic Systematics (Alfaro et al, 2003; Erixon et al, 2003; Wheeler & Pickett, 2008).


A problem in BI is the absence of convergence in the results when new data are adhered to data set. This parameter, the consistency (same groups recovered in different run and data set), is a adequate parameter to estimate the "efficiency" of a phylogenetic method. The "efficiency" is measured as the amount of nodes recovered to adhere more data in an statistical analysis.


Evenly, the a priori assumptions in BI (priors) has been debated because its subjectivity and inffluence in the results (Huelsenbeck et al, 2001; Rannala, 2002). The priors are parameters with probability and characteristics a priori. This assumptions are treated as random variables. So, the rate of substitution, the substitution model, rate of evolution, and the prior probability of the initial trees are asigned before the analysis.


In the common works, priors are estimated frequently using previous studies or subjectively. In the same way, is posible that the posteriori probability be influenced by the asumptions a priori, namely,the bayesian calculation is sensible to the estimate of priors.


Finally, the result of BI is a "topology" that show groupings of clades supposedly. Nevertheless, this “topologies” are not phylogenetic trees, but they are a representations of groups with "high probability" of to be sampled. So, the phylogenetic relationship are not recovered in the bayesian analysis. Too is interesting the frequentist vision of Bayesian Inference, where the more probable grouping of taxa is the correct. This approach is inappropriated because the hypothesis in Phylogenetics Systematics must be corroborated (in Popperian sense), a approach more adequate to scientific objective of Systematics.


Support


Some authors (Alfaro et al, 2003; Douady et al, 2003; Cummings et al, 2003) claims that the results of BI is equall to parametric bootstrapping (however, this statement is not very strong!). Nevertheless, others authors states that the BI has advantages on parametric bootstrap because its easy intepretation and speed. The parametric bootstrapping (one kind of support in phylogenetic systematics) generates new data from the initial topology and makes a new search using the new data and model. While the a posteriori probability of a bayesian topology is based in the more high frecuencies of the nodes recovered. So, the Bi is not similar to parametric bootstrapping, although yours results be equals in some studies.


Other methods of support as no parametric Bootstrap (Felsenstein, 1985) Bremer support (Bremer, 1994), Jackknife (1996), and Bremer support (Relative Bremer support) sensu Goloboff & Farris (2001), are considered approaches to estimate the support of nodes in a topology. However, an ideal measure of support must can estimate support using the evidence in favor and against of the resulting groups (nodes). So, i believe that the measure of support more adequate in phylogenetic Systematics is the Bremer support sensu Goloboff & Farris (2001). Furthermore, the measures of support that uses evidence in favor only can be considered as frecuentist, in the same way as BI.


Bibliography


  • Alfaro, M. E., Zoller, S., & Lutzoni, F. (2003) Bayes or bootstrap? A simulation study comparing the performance of Bayesian Markov chain Monte Carlo sampling and bootstrapping in assessing phylogenetic confidence. Mol. Biol. Evol., 20, 255-266.

  • Bremer, K. (1994) Branch support and tree stability. Cladistics, 10, 295-304.

  • Cummings, M. P., Handley, S. A., Myers, D. S., Reed, D. L., Rokas, A., & Winka, K. (2003) Comparing Bootstrap and Posterior Probability Values in the Four-Taxon Case. Systematic Biology, 52, 477-487.

  • Douady, C. J., Delsuc, F., Boucher, Y., Doolittle, W. F., & Douzery, E. J. P. Comparison of Bayesian and Maximum Likelihood Bootstrap Measures of Phylogenetic Reliability. Mol. Biol. Evol., 20, 248-254.

  • Erixon, P., Svennblad, B., Britton, T., & Oxelman, B. (2003) Reliability of Bayesian posterior probabilities and bootstrap frequencies in phylogenetics. Systematic Biology, 52, 665-673.

  • Farris, J. S., Albert, V. A., Kallersjo, M., Lipscomb, D., & Kluge, A. G. (1996) PARSIMONY JACKKNIFING OUTPERFORMS NEIGHBOR-JOINING. Cladistics, 12, 99-124.

  • Goloboff, P. A., & Farris, J. S. (2001) Methods for quick consensus estimation. Cladistics, 17, 26-34.

  • Huelsenbeck, J. P., Ronquist, F., Nielsen, R., & Bollback, J. P. (2001) Bayesian inference of phylogeny and its impact on evolutionary biology. Science, 294, 2310–2314.

  • Rannala, B. (2002) Identifiability of parameters in MCMC Bayesian inference of phylogeny. Systematic Biology, 51, 754-760.

  • Rannala, B., & Yang, Z. (1996) Probability distribution of molecular evolutionary trees: A new method of phylogenetic inference.

  • Wheeler, W. C., & Pickett, K. M. (2008) Topology-Bayes versus Clade-Bayes in Phylogenetic Analysis. Mol. Biol. Evol., 25, 447-453.

  • Yang, Z., & Rannala, B. (1997) Bayesian phylogenetic inference using DNA sequences: a Markov chain Monte Carlo Method. Mol. Biol. Evol., 14, 717-724.

domingo, 25 de noviembre de 2007

Homology

In cladistic analysis, the inference of homology has been previously suggested to be at least a two-step procedure: the first step is the hypothesis of correspondence of constituent features between two or more organisms. The second step subject these character hypotheses to the test of congruence (Rieppel & Kearney, 2002). In this sense it is not just the method used for phylogenetic inference which determines the quality of relationship hypotheses. Definition and selection of characters constitute a fundamental step (in the relationship hypotheses), because with them the crucial test in phylogenetic analysis is done.

A character is a logical relation established between intrinsic attributes of two or more organisms that is rooted in observation (Rieppel, 1988). So, “a meaningful character is thus based upon a character description that can in itself be evaluated... and potentially rejected” (Riepperl & Kearney, 2002). In morphological data, there are some classical characteristics that a good character must have: topology, connectivity, and the establishment of a one-to-one relationship of the parts being compared (Rieppel & Kearney, 2002). Although, there are other attributes as function and “special similarity” that are relevant in character definition (Rieppel & Kearney, 2002; Agnarsson & Coddington, 2007). Sometimes, topology correspondence, function and special similarity could conflict and quantitative methods for choosing among different criteria are necessary (Agnarsson & Coddington, 2007).

Other problem that morphological data face is coding. When we compare features in different organisms in some way, we are assuming some correspondence (not necessary topological). So, a presence-absence coding (See Pleijel, 1999) is telling nothing about that feature correspondence, and nothing about the taxa relationship. Becasuse, what forms the evidence in a cladistic analysis is the change among the character states, not the existence of different states (Brower, 2000).

In molecular data, the problem of character definition and character coding is different. Because it is widely accepted that similarity is equivalent to homology, but as those data also have homoplasy, similarity could not be seen as homology. Therefore, a criterion for defining molecular character hypotheses is necessary (de Pinna, 1991). The characters definition with DNA (or other molecular data) could be seen as a previous step, aligning with an algorithm. Or could be seen as a direct search of the optimal trees via direct optimization of the data with which the character hypotheses change during the search (Wheeler, 1996). The second approach is preferable, while it is testing the topology directly and is exploring different possibilities of the data, not just an alignment.

Finally, congruence is the test that corroborates the character as synapomorphies. The tool that allow to asses hypotheses of relationship, and concomitantly of character evolution, is parsimony. Parsimony maximizes congruence between the characters, so it maximizes the propositions of homology (Farris, 1983; Kluge, 1997; Sober, 1998).

Agnarsson, I. & Coddington, J. A. 2007. Quantitative tests of primary homology. Cladistics.
Brower, A. 2000. Homology and the inference of systematic relationships: some historical and philosophical perspectives. In Homology and systematics, coding characters for phylogenetic analysis (eds. Scotland, R & Pennington, R. T.). Taylor & Francis.
de Pinna, M. C. C. 1991. Concepts and tests of homology in the cladistic paradigm. Cladistics.
Kluge, A. 1997. Testability and the refutation and corroboration of cladistic hypotheses. Cladistics.
Pleijel, F. 1995. On character coding for phylogenetic reconstruction. Cladistics.
Rieppel, O. & Kearney, M. 2002. Similarity. Biological Journal of the Linnean Society.
Wheeler, W. C. 1996. Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics.
Farris SJ. 1983. The logical basis of phylogenetic analysis. In: Platnick, NI, Funk, VA, eds. Advances in Cladistics, Vol. 2. New York: Columbia University Press.

Homology

Homology is correspondence due to their shared ancestry. Homology assessment is a crucial step in phylogenetics analyses regardless of the type of data employed, since hypotheses of homology relate observations among taxa.

Every proposition of homology involves two stages which are associated with is generation and testing: the primary homology statement is conjectural based on similarity. The secondary level of homology is the outcome of a patter detecting analysis, the congruence test, and represents a test of the expectation that the observable match of similarities is potentially part of a retrievable regularity indicative of a general pattern (de Pinna, 1991).

When dealing with molecular data similarity guides primary homology -as in morphological characters- but there have been a tendency to believe that higher the similarity, the more likely that the sequences are homologous (Salemi and Vandamme, 2003; Patterson, 1988) where the fundamental homology statements are made at the level of the individual nucleotide bases. Conversely, the fundamental homology statement can be consider at the level of the sequence itself because the contiguous sequences are the homologous units that transform at prescribed costs among various states (Wheeler, 1999). Therefore Sequences themselves are treated as the fundamental units of homology and homology assessment is not consider a matter of aligning sequences and counting matches between them. Following “Direct Optimization’’ we can vary dynamically the primary homology hypotheses during tree search consequently the resulting hypotheses of homology are tested in conjunction with character congruence through parsimony (Wheeler, 2006). The testing of all characters against one another simultaneously constitutes the most severe test of congruence; In a phylogenetic context, the best alignment is the one that generates the most parsimonious tree when analyzed in conjunction with all relevant data (Philips, 2006).

The homology assessment have been shown to be influenced by the way we coding and the different interpretations of the criteria used to asses those statements demonstrating that the organismal variation is often conceptualized as characters and characters states in different ways (Scotland and Pennington, 2000). Even so, it is possible to have greater explicitness in the delimitation of morphological characters by detailed observation and the topologic criterion used, used in conjunction with special quality, and intermediate conditions of form (Rieppel and Kearney, 2002). In general, better guidelines concerning character conceptualization are required to help solving this.

Similarity and conjunction have proposed as tests for homology but only congruence serves to this respect. Similarity, as mentioned above, guides the assessment of the homology conjecture. Conjunction is an indicator of non-homology, but it is not specific about the pair wise comparison where non-homology is present, and depends on a specific scheme of relationship in order to refute a hypothesis of homology (de Pinna, 1991).

On Homology

Introduction
The importance of homology has been discused by several authors (Patterson, 1988; Wagner, 1989; de Pinna, 1991; Rieppel & Kearney, 2002; Agnarsson & Coddington, 2007). So, homology is a crucial basis in the Systematics. The most simple meaning of homology is equivalence of parts (de Pinna, 1991). In 1982, Patterson states that homology is equal to synapomorphy. So, the synapomorphic characters must be homologous. Nevertheless, the symplesiomorphic characters could be homologous (homologies at a higher level).
There are two kinds of homologies; the primary homology - conjectures or hypothesis about common origin of characters -, and the secondary homology - the tested hypothesis – (de Pinna, 1991).

Pattern or process
A question about homology is the significance of evolutionary process in the identification of homologous characters. Implicitly, the evolutionary events are bounded to the analysis of characters (Lee, 2002). However, Brower (2000) stated that “the similarities - homologies - between taxa represent the only necessary ontological foundation for the construction of cladograms and hypotheses of taxonomic grouping”. So, the evolutionary asumptions are not necessary in the identification of homologies; but the homologous characters can be explained by evolutionary process (Rieppel & Kearney, 2002).
Equally, some authors claims that the phenomenon of circularity in homology (need of a priori topology) is undesirable because the recognition of homologous characters is conditionated to mapping of them in a initial topology. Nevertheless, the circularity is not a “big” problem, because the tested homologies (mapped hypothesis of homologies) are “tested hypothesis” attached to new set of analysis (new tests).

Coding
An inherent point in the identification of homologous characters is the character's coding. A inadequate definition of characters produces bias in the identification of homology. Some author claims that the real problem in homology is the character's coding. So, for example, there is the belief that the character's coding is linked to knowledge of researcher about taxon. So, the “eyes” of an experienced researcher would discriminate and describe “best” characters and character's states that an non-experienced researcher.
A approach used in the the character's coding was the morphometric analysis (biometry). However, the biometry is not useful in homology because the complexity of the biological estrucures and its incompatibility with the statistical multivariate analysis (Bookstein, 1994).

Tests
The three tests of homology (similarity, conjuction, and congruence) are secuencial in the identification of homology. Nevertheless, the similarity is not a test (in Popperian sense), similarity is a conjecture of homologous characters - primary homology – (de Pinna, 1991) . The conjuction and congruence (agreement in supporting the same phylogenetics relationships) are the “hard” tests of homology – the secondary homology from de Pinna - (Rieppel and Kearney, 2002). Although there is interdependece among them (tests), it not means that the three tests will be one only. So, the result is a hypothesis of homology corroborated, but they is not definitive (“true homology”).

Molecular homology
The identification of homology in molecular characters (nucleotide sequences) presents problems not found in other kinds of character data. For example, although each base position presents one of four identical states (A, C, G or T), the number of these positions is likely to vary, that is homologous nucleotide sequences may differ in length (Wheeler, 1996). Further, Patterson (1988) claims that the tests of homology in molecular characters are equal to morphological characters. However, the significance of tests (similarity, conjuction, and congruence) is different because the similarity is the most crucial test, while in morphological characters is test of congruence.
A point of discussion in the identification of homologies in molecular sequences is the need of an alignment to determine sites or homologous fragment. However, this way is considered inadequate because alignment is generated using a priori costs and asumptions. Wheeler (2003) proposes a synapomorphic-based alignment methods - Implied alignment – (IA) that identifies homologies in the topology using Direct Optimization (DO). So, Implied alignment generating all posibles alignments and all posibles homologies are analized. This method may be efficient to identify homologies. Nevertheless, the a priori asumptions are inherent to alignments.

Bibliography

  • Agnarsson, I., & Coddington, J. A. (2007). Quantitative tests of primary homology. Cladistics, 23, 1-11.
  • Brower, A. V. Z. (2000). Evolution is not an assumption of cladistics. Cladistics, 16, 143–154.
  • de Pinna, M. C. C. (1991) Concepts and tests of homology in the cladistic paradigm. Cladistics, 7, 367-394.
  • Lee, M. S. Y. (2002). Divergent evolution, hierarchy, and cladistics. Zoologica Scripta, 31, 217–219.
  • Patterson, C. (1988). Homology in classical and molecular biology. Molecular Biology and Evolution, 5, 603-625.
  • Rieppel, O., & Kearney, M. (2002) Similarity. Biological Journal of the Linnean Society, 75, 59-82.
  • Wagner, G. P. (1989) The biological homology concept. Annual Reviews of Ecology and Systematics. 20, 51-69.
  • Wheeler, W. C. (1996). Optimization alignment: the end of multiple sequence alignment in phylogenetics?. Cladistics, 12, 1-9.
  • Wheeler, W. C. (2003) Implied alignment: a synapomorphic-based multiple alignment method and its use in cladogram search. Cladistics, 19, 261-268.

lunes, 29 de octubre de 2007

On Evidence

Evidence are a group of events that support or not a hypotesis. A pure observation not must be considered as evidence, a pure observation (observations are not bounded to theorical basis) is not evidence. The evidence's quality is related to hypotesis which it is bounded. So, all evidence is not relevant for a hypotesis. Evidence is not considered as “good” or “bad”, simply it is support a event rather other. The pure observations (without theoric basis) is considered “good” or “bad. So, in the historical sciences the observation of event is not enough, these observation must be bounded a theorical background. The “smoking gun” is a example which the evidence is obtained, but all the observations in the event ( murder) must be related to a theory. The observations that are not related to murder, it are not evidence.