Content of review 1, reviewed on December 06, 2018

The authors propose a new biclustering method and demonstrate its applicability in a set of breast invasive carcinoma cohorts. The method stems from the idea of gene proximity based on mutual agreement in the sets of samples where genes from a gene pair are expressed highest or lowest. Based on proximity, a gene graph is constructed, biclusters correspond to densely connected genes and samples overrepresented on the subgraph edges.

The idea of gene similarity defined by the size and significance of the overlaps between the sample sets that fall into percentile sets of the inidvidual genes is plausible and appealing.

First, I have a conceptual comment to the algorithm design. The construction of biclusters completely ignores samples. In other words, the authors make an implicit assumption that genes appearing in a densely connected subgraph get co-expressed in a similar set of samples. In general, this does not have to be satisfied and the proposed algorithm degenerates to regular clustering then. As an example, take a triplet of genes g1, g2 and g3, where all the gene pairs are connected by an edge. g1 and g2 match in s1 and s2 (these samples make the overlap), g2 and g3 match in s3 and s4, g1 and g3 match in s5 and s6. In TuBA, this subgraph is of the same quality as a subgraph of g4, g5 and g6, where all the gene pairs match in s7 and s8. In any reasonable definition of biclustering, the second case should strongly be preferred. Biclustering is defined as CONCURRENT clustering in both the dimensions which is not met in TuBA and TuBA seems to be inferior to many existing biclustering algorithms from this point of view. On top of that, the authors tend to interpret their results mostly in terms of genes and work with non-overlapping patient cohorts which makes any comparison in this dimension difficult. These features lead me to the conclusion that the algorithm should rather be presented as a regular clustering with a new gene proximity measure. Related work as well as the comparative experimental evaluation should be changed accordingly. To be more specific, I would appreciate a comparative evaluation of the newly proposed proximity measure. The benchmarks could be global criteria (for example, Spearman rank correlation mentioned in the paper) as well as local ones (if they exist).

Second, there are some principal limitations in the algorithm design. As far as I understand TuBA algorithm, it does not allow for arbitrarily overlapping biclusters, or I should rather say gene clusters. The overlaps are limited to the restricted areas that make close neighborhood of the individual graph cliques. This is a reasonable limitation when working with multifunctional genes and this feature should be touched in discussion at least. Further, the algorithm includes the largest clique identification problem which is computationally difficult. Runtime of the algorithm should be studied and discussed in more detail. For the moment, the authors only recommend a heuristic constraint (200,000 edges) which does not fully take into consideration the true complexity of the largest clique identification problem. Runtime should be considered as a criterion in terms of the concluding comparative evaluation.

Presentation and language are very good. I identified one inconsistency in Figure 3 and its accompanying description. The authors claim that the individual significance curves are as follows: (i) top 20% (blue), (ii) top 10% (red), and (iii) top 5% (green). I guess that it is just the opposite (5% blue, 20% red), the overlap significance and sensitivity should contradict.

The description of the graph-based algorithm in Step 2 (line 151-154) is unclear. What happens if having more candidate largest cliques even after the union of overlapping ones?

Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below. I declare that I have no competing interests.

I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published.
I agree to the open peer review policy of the journal.

Authors' responses to reviews: (https://drive.google.com/open?id=1Opuc4BLagbgGxigy4sSyI-c7AeX35BTQ)

Source

    © 2018 the Reviewer (CC BY 4.0).

Content of review 2, reviewed on April 02, 2019

The authors propose a new biclustering method and demonstrate its applicability in a set of breast invasive carcinoma cohorts. The method stems from the idea of gene proximity based on mutual agreement in the sets of samples where genes from a gene pair are expressed highest or lowest. Based on proximity, a gene graph is constructed, biclusters correspond to densely connected genes and samples overrepresented on the subgraph edges. The idea of gene similarity defined by the size and significance of the overlaps between the sample sets that fall into percentile sets of the inidvidual genes is plausible and appealing.

This manuscript is a resubmission of the previous paper that I also reviewed. I checked the changes that the authors made and verified how the improvements answer my earlier comments:

  1. I had a conceptual comment to the algorithm design. I proposed to present the algorithm rather as a regular clustering with a new gene proximity measure. The authors keep their algorithm in the class of biclustering methods, however, they introduced a new experimental section named Enrichment of bicluster samples in top (bottom) sample sets of bicluster genes in which they showed that their algorithm tends to generate biclusters with core functional signatures, i.e., it is possible to identify core sample that comprise the bicluster and whose occurrence is statistically significantly enriched in the top and bottom sample sets of nearly each gene in the bicluster. This experiment sufficiently answers my doubt that genes appearing in a densely connected subgraph do not have to be co-expressed in a similar set of samples.

  2. There are some principal limitations in the algorithm design. The algorithm includes the largest clique identification problem which is computationally difficult. I proposed to study runtime of the algorithm in a more detail. The authors did so in a new section named Runtime analysis. The section confirmed the expectation that TuBA has longer runtime than most of the other biclustering algorithms. Despite of the inconvenient asymptotic bounds, the authors demonstrated that the algorithm can be applied to current gene expression datasets with tens of thousands genes and thousands of samples. The runtime can be influenced by the setting of the significance level of overplap between percentile sets

  3. The authors answered a few detailed technical comments that referred to Figure 2 and the Graph based algorithm.

Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below. I declare that I have no competing interests.

I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published.
I agree to the open peer review policy of the journal.

Authors' responses to reviews:

Reviewer 3: In this paper the authors present a tunable biclustering algorithm (TuBA), based in the use of a proximity measure. TuBA has been applied to Breast Cancer Analysis. The paper is well organized and the contents include an extensive analysis on different datasets. Nevertheless, I have the following concerns regarding its appropriateness for being accepted.

Response: We thank the reviewer for their positive comments and hope that in the revised manuscript, we have clarified and addressed these concerns.

Reviewer 3: In section METHODS, subsection "Proximity Measure" does not provide enough details in order to reproduce its computation. It is not easy to understand how it does evaluate biclusters.

Response: TuBA’s proximity measure is based on computing the significance of the overlap between the samples that exhibit high (or low) expression of any given gene pair. As noted in lines 110-114, “The number of samples shared between any pair of percentile sets follows the hypergeometric distribution; therefore, we can compute the significance (p-value) of overlaps between pairs of percentile sets based on the numbers of shared samples by using the one-sided Fisher's exact test.” The lower the p-value of overlap, the stronger is the association between the gene-pair. We illustrate the basic idea and the setup of the test underlying our proximity measure in Figure 1.

Reviewer 3: In section METHODS, subsection "Datasets" it is not clear how many datasets have been finally analysed. Although in line 248 it is said that "we applied TuBA to three BRCA datasets", TCGA, METABRIC and GEO, from line 258 some other datasets are introduced: Bicmix or GTEx.

Response: The primary and clinically relevant results of the paper are based on the application of TuBA to three large breast cancer (BRCA) datasets – TCGA, METABRIC, and GEO. We additionally used the Genotype-Tissue Expression (GTEx) and ESTIMATE datasets to reaffirm that some of our biclusters from the three BRCA datasets were associated with normal/stromal tissue gene expression signatures. Moreover, for comparing TuBA with the Bicmix algorithm, we applied TuBA to the breastCancerNKI dataset. We have now separated this subsection in Methods to clearly indicate primary and validation datasets (lines 248-289).

Reviewer 3: In section RESULTS, subsection "Benchmarking TuBA's Proximity Measure", proximity measure has been checked against correlation coefficients such as Pearson or Spearman, but not against other evaluation measures specific for biclusters.

Response: TuBA’s proximity measure, like Pearson’s and Spearman’s correlation coefficients, is a pairwise measure of proximity. It enables us to identify gene-pairs that show strong association by virtue of a significant overlap in their top (or bottom) samples. Such gene-pairs are represented graphically as a pair of nodes linked together; the link represents the set of samples that are shared in their respective top (or bottom) percentile sets. Thus, our proximity measure provides the basis for the generation of graphs that are subsequently analyzed by our iterative graph-based approach to identify biclusters. By itself, the proximity measure does not evaluate biclusters.

We should note that TuBA does not involve modeling of gene expression data in order to identify biclusters. Thus, bicluster evaluation measures based on models of patterns of gene expression data are not applicable to its results. However, in the Enrichment of bicluster samples in top (bottom) sample sets of bicluster genes subsection of Results, we are now proposing a simple metric to assess the overall quality of the biclusters discovered by TuBA (Lines 404-412): “… the FDR values for each gene (sample) in any given bicluster can be viewed as their scores – the closer the value of the FDR is to zero for a gene (sample), the stronger is the association of the gene (sample) to the bicluster. We used the FDR values for the genes and samples within bicluster i to evaluate its overall quality, Q(Bi), defined as the minimum of the fraction of the genes in the bicluster with FDR < 0.05 or the fraction of samples in the bicluster with FDR < 0.05. Q(Bi) takes values between 0 and 1, where values close to zero indicate weak associations between the constituent genes and samples within the bicluster, while values close to one would indicate strong associations.” Supplementary Table 5 is now updated to include these quality scores for the biclusters.

Reviewer 3: In section RESULTS, subsection "Choice of parameters for TuBA" it not specified a clear strategy for choosing the parameter values when analysing a new dataset.

Response: We have discussed the considerations that underlie the choices for TuBA’s parameters in the Tuning TuBA subsection of Methods, also illustrated in Figure 3. For the choice of the second parameter (Overlap significance cutoff), we proposed the following heuristic in the same section: “the cutoff for the significance level of overlap should be such that a decrease in the significance level by an order of magnitude leads to an 40-60% increase in the number of edges that get added to the graph.” (Lines 233-237)

Reviewer 3: Furthermore, being the results section so large, I also suggest to include a summary table or diagram in order to provide the reader with a quick snapshot of the results.

Response: We have summarized the key results and discussions in Supplementary Figure 10.

Source

    © 2019 the Reviewer (CC BY 4.0).