Skip to main content

How Search Works in Keenious

Keenious turns your plain-language question into a structured academic search and ranks results transparently using both meaning and scholarly cues.

Overview

Keenious helps students and researchers discover and understand academic literature. This page describes how its search works: how a query is matched against a curated index of ~188 million publications spanning all disciplines and more than 100 languages, and what determines the order of results. The same search engine runs whether you type a query in Search or the AI searches on your behalf in Chat. How the index is built is documented in How Keenious Curates Its Academic Search Index.

How It Differs from a Traditional Search

A Keenious search is made for overview as much as for individual results. Instead of one long ranked list to sift through line by line, it returns a set of the best-matching publications and analyses that set as a whole. The results come grouped into research areas: the strongest matches and a picture of the directions the literature takes, at the same time. Digging deeper happens by opening the research area that interests you, rather than paging further down a list.

The matching is semantic: a query is understood by its meaning, whatever its phrasing. Exact requirements can be added on top with Boolean-style operators (quoted phrases, AND, OR, and NOT) and filters. The same search works as a loose description of a topic or as a tightly constrained query, depending on the task.

Search matches against titles and abstracts, not full text. A publication can discuss a topic in its body without it being mentioned in the title or abstract; such a publication will not match a query on that topic.

Matching

A query is matched in two ways at once. By meaning: every title and abstract in the index is stored as an embedding (a representation of its meaning), and the publications closest in meaning to the query are found, whatever words they use. teenage mental health and social media matches publications about "adolescent well-being and problematic smartphone use". By keyword (BM25, short for "Best Matching 25", a standard keyword-ranking method): the same query's exact words count too, weighted toward distinctive terms (gene names, acronyms, place names) that meaning alone can treat loosely.

Publications as points in a space of meaning: the query sits among its closest publications, which are retrieved regardless of their wording, while unrelated publications sit far away

A publication that scores well on both meaning and keywords ranks highest, with meaning carrying more weight (the merge is a rank fusion). Quoted phrases and filters sit on top of all of this as hard requirements: see Search Syntax.

Ranking

Relevance is how well a publication matches your query: its combined score from the meaning and keyword matching above. Results are ordered by relevance first. Scholarly signals then adjust the order, each offering a modest adjustment that reorders similarly relevant publications without overriding relevance. The signals apply independently and are listed in no particular order:

  • Citations: more-cited publications rank higher. Citation counts are computed within the curated index, so they can be lower than in databases that count against a broader corpus.

  • Recency: recent publications receive a small boost that fades with age.

  • Field-weighted citation impact (FWCI): citation performance normalized by field, year, and document type, so publications from low-citation and high-citation fields are compared on the same scale.

  • Peer-review status: venues listed in the Norwegian Scientific Index receive a boost, Level 2 channels more than Level 1. Venues not listed receive no boost, but no penalty either.

  • Venue and publication type: journal and conference publications rank slightly above works without an identifiable venue; review and research articles rank slightly above other document types.

  • Language match: publications written in the query's language are boosted. English-language work dominates global citation counts, so without this signal, queries in other languages would return mostly English results.

Scope (Relevance Floor)

For an in-depth guide to how Scope interacts with candidate pools, safeguards sorting, and shapes entity discovery, see Controlling Search Scope and Relevance Floors.

In traditional keyword databases, records either match or do not. Semantic matching has no such binary boundary: relevance exists along a continuous spectrum, and any single cutoff would be arbitrary. To give you direct control over this boundary, Keenious provides Scope, a relevance floor setting how closely a publication must match your query to appear in results.

Concentric relevance boundaries in meaning space around the search query, with the corresponding 4-segment meter.

Scope offers four levels, represented in the interface by a segmented meter:

Level

Meter

Relevance Floor

Best Used For

Focused

1 segment

High semantic similarity threshold across publications.

Pinpointing core, highly specialized literature on a specific question.

Balanced

2 segments

Calibrated relevance floor balancing precision with topical breadth. This is the default setting.

Standard topic exploration and general research surveys.

Exploratory

3 segments

Relaxed relevance threshold capturing peripheral connections.

Surfacing interdisciplinary connections, adjacent concepts, and emerging themes.

All

4 segments

Removes the similarity floor completely.

Broad discovery returning top candidates regardless of semantic distance.

The Scope menu displays the exact number of matching publications available at each tier, allowing you to see how many papers meet each relevance threshold before switching. Scope requires search text to compute semantic similarity; during filter-only browsing without keywords, Scope is inactive.

The Result Set

A search retrieves candidate publications up to a selected pool ceiling: 300 publications by default, adjustable through the results count menu up to 1,000 (or up to 10,000 when browsing with filters alone).

When Scope is active, publications falling below the selected relevance floor are excluded. The result count therefore indicates "Up to 300" (or your selected size). If fewer publications meet the relevance threshold than the pool ceiling, the returned count represents an exact total of qualifying papers.

Everything that follows operates on this scoped set:

  • Research areas are computed exclusively from the returned publications.

  • Sorting reorders the publications within the scoped set. When you sort by citations or publication date, you view the most cited or newest papers among those that satisfy the relevance floor. Scope makes sorting dependable: without a relevance floor, sorting by citations would surface landmark papers from adjacent fields that only distantly relate to your query.

Systematic reviews and legal discovery requiring documented, exhaustive retrieval should use comprehensive bibliographic databases alongside Keenious. Keenious is built for literature exploration and topic synthesis rather than exhaustive records of legal record.

Every search is saved with a permanent link, and opening that link always shows the same results. Typing the same query again later runs a fresh search against the updated index.

Sorting

Results are ordered by relevance by default (the ranking described above). The Sort control reorders the retrieved set:

  • Relevance (Best match): Orders by combined semantic and keyword matching, adjusted by scholarly signals.

  • Most cited: Highest citation count in the curated index first.

  • Newest: Most recently published first.

Sorting reorders the publications already in the scoped result set; it does not pull in different papers from the wider index.

Research Areas

The results of a search come grouped into research areas: clusters of related publications within the result set. Publications whose meanings sit close together (by the same embeddings used for matching) form an area, and each publication belongs to exactly one. The number of areas scales with the size of the result set, up to 15.

Result publications shown as points, colored into three labeled clusters: each cluster of nearby publications is a research area

Each area is named by a language model for what sets it apart from the rest of the results, rather than for the query: a search about social media and teenage mental health gets "Cyberbullying and Online Harassment" and "Screen-Time Interventions", not "Social Media Studies". The names are not a fixed taxonomy, and the same publication can sit under a differently named area in another search.

Selecting research areas filters the result list; it does not change the ranking within them. The areas are computed from the result set itself, so the same search always produces the same areas, and changing the search (the query, the filters, the scope, or the result size) recomputes them.

Search Syntax

Example

Effect

gene editing therapy

Matched by meaning and keywords (default)

"CRISPR-Cas9"

Must appear as an exact phrase in the title or abstract

arctic ecosystem AND climate change

Each side's words must all appear, in any order

"CRISPR" OR "TALEN"

At least one must appear

-mice, NOT mice, or -"in vitro"

Must not appear in the title or abstract

2021 (a bare four-digit year)

Publications published that year are boosted; others still appear

Quoted phrases. A quoted phrase is a hard requirement: every result must contain it in its title or abstract. A multi-word phrase must appear as consecutive words in the given order: "gene therapy" matches "a gene therapy trial" but not a publication where gene and therapy only appear in separate sentences. Matching ignores capitalization, accents ("Zurich" matches "Zürich"), and punctuation ("CRISPR-Cas9" also matches "CRISPR Cas9"). It does not ignore word forms: quoted matching has no stemming, so "vaccine" does not match a publication that only writes "vaccines", and "mouse" does not match "mice". Unquoted text is unaffected by this: keyword matching on unquoted words handles word forms, and semantic matching is independent of wording altogether.

A quoted phrase also remains part of the query for semantic and keyword matching; the quotes add the requirement on top rather than replacing the term's role in matching.

AND segments. Uppercase AND splits the query into segments that must each appear in the title or abstract: a bare segment requires all of its words, in any order; a quoted segment requires the exact phrase. arctic ecosystem AND "climate change" requires arctic and ecosystem somewhere in those fields, plus the exact phrase climate change. Several quoted phrases in a row are an implicit AND. Lowercase and is an ordinary word. Operator tokens are not themselves matched text: X AND Y and X Y match by meaning identically and differ only in the requirement.

OR groups. OR operates between adjacent terms, quoted or bare: in "CRISPR" OR "TALEN" "off-target", at least one of CRISPR or TALEN must appear, and off-target must appear; teenagers OR adolescents works the same without quotes. Longer chains work the same way ("mouse" OR "mice" OR "murine"). OR binds tighter than AND: ecosystem AND "climate change" OR warming requires ecosystem plus at least one of the other two. Lowercase or is an ordinary word.

Exclusions. -term and -"phrase" remove every publication whose title or abstract contains the term (NOT is an uppercase alias: NOT mice and NOT "in vitro" mean the same as -mice and -"in vitro"). Excluded terms follow the same rules as quoting: capitalization, accents, and punctuation are ignored, but word forms are not, so -mouse does not remove publications that only write "mice". Excluded terms are stripped from the query before matching: they only remove results, they do not influence what the rest of the query matches.

Years. A bare four-digit year (1900 up to next year) boosts publications published in that year; publications from other years still appear, without the boost. Several years can be given (2023 2024), and the year also remains part of the ordinary query text. A hard cutoff is a job for the year filter, not the query.

Operators combine freely: gene editing "CRISPR-Cas9" -mice matches the concept gene editing semantically, requires CRISPR-Cas9 in the title or abstract, and removes publications mentioning mice. All quoted and excluded terms are matched against titles and abstracts only: a quoted term excludes any publication that does not use it in those two fields, even if the term appears in the publication's full text.

Filters (publication year, document type, peer-review status, open access, and others) are hard constraints rather than ranking signals: a filtered-out publication is removed before ranking and will not appear regardless of how well it matches.

Example Searches

Query

How It Works

microplastics in arctic marine food webs

Plain language, matched by meaning: the wording need not match the publications

"perovskite solar cells" stability 2024

Requires the exact phrase, matches stability semantically, boosts 2024

"CRISPR-Cas9" AND "off-target effects" -cancer

Requires both phrases, excludes publications mentioning cancer

"CRISPR-Cas9" "off-target effects" -cancer

Same as above: adjacent quoted phrases are an implicit AND

"machine learning" AND "drug discovery" OR "drug design"

Requires machine learning, plus either drug discovery or drug design

"carbon capture" "direct air capture"

Title or abstract contains both phrases: more precise, but it misses those using one term

carbon capture direct air capture

Matches either or both semantically: broader, surfaces publications connecting the two

Why a Publication Does or Doesn't Appear

A result looks off-topic. Semantic matching retrieves by closeness of meaning, and the matched concept is normally visible in the result's title or abstract. Quoting a term that must appear, or excluding one that shouldn't, narrows the results.

An expected publication is missing. Common causes:

  • The wording only appears in the full text. Titles and abstracts are the searchable fields; a topic that is not visible there does not match.

  • A quoted or required term is not in the title or abstract. Quoted phrases and AND parts are hard requirements. Without them, the publication can still match semantically.

  • A filter excludes it. A year range or peer-review filter removes everything outside it.

  • It is excluded by the active Scope. The publication connects to your topic, but falls below the semantic threshold set by Focused or Balanced. Broadening the scope to Exploratory or All includes it.

  • It is not in the index. Book chapters, master's theses, meeting abstracts, retracted works, and several source types are excluded from the OpenAlex dataset.

  • It is outside the result set. A search returns a bounded pool of best-ranked publications (see The Result Set). A publication can match without making the cut: a more specific query, an expanded Scope, or a larger result size brings it in.

  • It is very new. The index synchronizes with OpenAlex regularly; a publication from the last few days may not be indexed yet.

Frequently Asked Questions

Does Keenious search the full text of publications? No: titles and abstracts only. A publication whose topic is visible only in its body will not match a query on that topic.

Does it matter whether I write and or AND? Yes. In capitals, AND is an operator that splits the query into required segments. In lower case, and is an ordinary word matched by meaning.

Can the AI make up a publication? No. Every result is a publication from the index. AI is used to match queries by meaning and to name research areas, never to generate results.

Why don't citation counts match Scopus, Web of Science, or Google Scholar? They are computed within the curated index: only citations from publications that are themselves in Keenious count. The numbers run lower but are internally consistent (see How Keenious Curates Its Academic Search Index).

Why did I get fewer results than the size I chose?
The active Scope floor, quoted terms, or active filters excluded publications that did not meet the criteria. In that case, the number shown is an exact count of qualifying matches. See Scope (Relevance Floor) and The Result Set.

If I run the same search next month, will I get the same results? Opening a saved search by its link always shows the same results. Typing the query again is a new search, and it can differ in results as the index is updated with new and corrected records.

Did this answer your question?