Text & data mining
This page is for publishers and rights-holders. It sets out how the Evidence TAP research project at the University of Cambridge intends to use text and data mining (TDM) of the research literature, and the safeguards we apply.
The project
Evidence TAP (the Cambridge Traceable AI Pipeline) is a non-commercial academic research project, hosted at the University of Cambridge departments of Computer Science, Zoology and Education. Its goal is to create living evidence databases for particular fields, beginning with conservation and education, with others to follow. A living evidence database continuously ingests the published and grey literature, screens it for relevance, appraises study design, and extracts structured findings, so that systematic reviews stay current for evidence-based policymaking rather than going out of date the moment they are published.
Delivering this depends on analysing the full text of research articles. We are committed to doing so responsibly, transparently, and in a way that respects publishers' infrastructure and licensing.
Frequently asked questions
- What is the purpose of the text and data mining?
- Non-commercial academic research. We build living evidence databases for particular fields, beginning with conservation and education, that keep systematic reviews current for policymakers. Mining full text lets us screen studies for relevance, appraise their design, and extract structured findings that remain traceable to their source. Read the pipeline paper → Text and data mining is enabled for non-commercial research by the copyright exception.
- Is the research team using any third-party software or tools to do this mining?
- All of our software is built in house for crawling, and is designed to be as considerate as possible of third-party API rate limits and keys. We do not use commercial scraping services or resell access.
- Can the team specify what tools they are using for this exercise?
- We use a custom software stack written in OCaml and Python that performs local-only analysis of full text. It uses a locally deployed instance of GROBID, local mirrors of CORE (core.ac.uk), Crossref and PubMed Central (PMC), and locally deployed large language models to perform a variety of metadata analysis. No article content is sent to any external or third-party service.
- Where does the copying and analysis take place?
- Entirely in the United Kingdom. All copying, storage and analysis takes place on University of Cambridge servers hosted in the Department of Computer Science and Technology (the Computer Laboratory). No article data is processed or stored outside the UK.
- How can we identify or allowlist your crawler?
- Our crawler requests identify themselves with a descriptive User-Agent string, EvidenceTAP/1.0 (+https://evidencetap.org/text-and-data-mining), and originate from University of Cambridge network ranges. If you would prefer we use a particular access route, respect a specific rate limit, or be added to an allowlist, please get in touch and we will adjust accordingly.
- Can you confirm the results will be closed and only available to Cambridge-authorised users?
- Yes. All downloaded data is stored on University of Cambridge servers that are firewalled to provide access only to the local research team and named collaborators. There is no public access to the underlying full-text corpus.
- Will the full text be redistributed or republished?
- No. The full-text corpus is never redistributed or republished; it remains firewalled on University of Cambridge servers. The living evidence database we produce contains derived outputs (structured findings and metadata, each linked back to the original publication), and these may be shared openly so that the evidence stays traceable to its source. We do not reproduce substantial portions of the original articles.
Get in touch
We are happy to answer questions, discuss access arrangements, or provide further detail about our infrastructure and safeguards. Please contact Professor Anil Madhavapeddy (avsm2@cam.ac.uk).