Consensus — AI Search That Answers Research Questions

Consensus answers a plain-language research question with findings extracted from papers and a meter showing how they divide. Excellent triage, unusable as evidence.

Official Site https://consensus.app
Category Literature
Pricing Freemium
Rating ★★★☆☆ (3/5)

What Consensus is

Consensus is an AI search engine for academic literature that takes a question in ordinary language — “does intermittent fasting improve insulin sensitivity?” — and returns papers with a one-line extracted finding from each, plus a summary meter showing how the retrieved studies divide between yes, no, and mixed.

It runs over the Semantic Scholar corpus, so its coverage is Semantic Scholar’s coverage. The genuinely novel part is the layer on top: a language model reads each paper’s claims and reduces them to a stated position on your question. That reduction is simultaneously why the tool is fast and why it must never be cited. Everything below follows from holding both of those facts at once.

Why researchers use it

  • You ask the actual question — no Boolean, no MeSH terms, no guessing the field’s preferred vocabulary before you know the field.
  • The meter shows contestation instantly — the fastest way to find out that a “well-established” finding is in fact split down the middle.
  • One-line findings per paper — enough to decide whether an abstract is worth opening, across twenty papers in the time it takes to read three.
  • Study-type filters — restrict to RCTs, systematic reviews, or by sample size, which sharpens the retrieved set considerably.
  • Links straight to sources — every extracted claim carries its paper, so verification is one click rather than a new search.

Where it fits in a research workflow

Consensus belongs to the first hour on an unfamiliar question, and to almost no hour after that. Its job is orientation: is there a literature here, does it agree, what vocabulary does it use, and which five papers should you actually read?

After that first hour it hands off. Real searching goes to Semantic Scholar, Scopus, or PubMed with terms Consensus helped you discover. Mapping the field goes to Connected Papers or Litmaps. Structured extraction across a defined paper set goes to Elicit. Checking whether a specific finding survived replication goes to scite, which reads citation intent rather than paper conclusions and is the better instrument for that question. Consensus is the doorway, not a room.

Getting started

Ten minutes, and the discipline matters more than the setup.

  1. Ask one specific, answerable question — a relationship between two named things. Broad topics return broad noise; “does X affect Y in Z population” returns something usable.
  2. Read the meter, then immediately ignore it and open the top five source papers. The meter’s only legitimate use is telling you whether to expect agreement.
  3. Apply the study-design filters. An unfiltered meter mixes a 12-person crossover study with a 40,000-participant cohort as if they were equal votes, because to the meter they are.
  4. The step people skip: check one extracted finding against the paper’s actual abstract. Do it early, do it once, and you will calibrate how much to trust the layer for the rest of your use.

Consensus vs the alternatives

AlternativeDoes it betterPick it if
ElicitStructured extraction of methods, populations, and outcomes into a tableYou need comparable data across a defined set of papers
sciteWhether later papers support or contradict a specific findingYou are asking whether a result replicated
Semantic ScholarReal search, free API, no extraction layer between you and the paperYou want the literature, not a summary of it
PerplexityGeneral web and news alongside academic sourcesYour question is not purely a research-literature question

Consensus is the fastest of these at telling you whether a question is contested. It is the weakest of them at telling you anything you could write down.

Cost, licensing, and your data

Freemium: a limited number of AI-summarised searches per month at no cost, with a subscription for unlimited use and the deeper synthesis features, plus student and institutional pricing.

Nothing sensitive leaves your side of the screen if you use it as intended — you type a question about published literature and receive public bibliographic data back. The consideration that does apply is subtler and worth naming: your queries describe your research direction, in detail, to a commercial third party. For most work that is unremarkable. For a competitive grant application or an unpublished hypothesis you would not discuss at a conference, a search log held by a vendor is a small but real disclosure, and the same caution applies to every AI search tool in this directory.

The honest review

Strengths. The consensus meter earns its place for one specific job: puncturing false certainty. Ask about something your field treats as settled and watch the meter come back split, and you have learned in fifteen seconds something that a week of reading in one direction would have hidden from you. As an instrument for detecting contestation, it is genuinely good.

Limitations. The extraction layer flattens exactly what matters. Effect size, population, dose, follow-up duration, study quality, and conflict of interest all vanish into a directional vote, and the meter weighs a pilot study and a definitive trial identically. That is not a bug that will be fixed; it is what reducing a paper to a position on a question necessarily does. Coverage inherits Semantic Scholar’s gaps, which are significant outside the sciences. And the interface’s confidence is out of proportion to its epistemics — a clean percentage next to a question invites a reader to treat it as a finding, which it is not.

Verdict. Adopt it as a fifteen-minute orientation tool at the start of an unfamiliar question. Skip it entirely for evidence synthesis, for anything going into a methods section, and for any question where the answer depends on population or dose. The condition that flips the answer is what you plan to do with the output: read further, useful; write it down, dangerous.

When NOT to use this Never cite the consensus meter, and never let it stand in for reading. A meter reading “78% yes” is a summary of what a language model extracted from whatever the search happened to retrieve — it is not a pooled estimate, it has no inclusion criteria, and it cannot be reproduced by a reader. If a claim matters enough to appear in your writing, it matters enough to trace to the paper and read the methods. This is the exact failure mode described in the AI research integrity guide.

Common questions

Is Consensus free?

There is a free tier with a monthly cap on AI-summarised searches, and a paid subscription for unlimited use plus the deeper synthesis features. Student and institutional rates exist.

Can I use Consensus for a systematic review?

No. It has no reproducible search strategy, no documented inclusion criteria, and no way for a reader to reproduce your result — three things a systematic review must have. Use it to discover terminology and seminal papers, then run and report a proper database search.

How accurate is the consensus meter?

It accurately reflects what a language model extracted from the papers the search retrieved, which is not the same as reflecting the evidence. It ignores study quality, sample size, population, and effect size, and it is sensitive to what the retrieval step happened to surface. Treat it as a signal of contestation, never as a result.