The sources are different: a researcher works through journal articles and preprints; a consultant needs government policy, regulation, industry reports, planning documents, company filings and specialist research; an analyst watches earnings calls, product releases, market data and regulatory decisions; and a journalist needs court filings, public records, interviews, company statements and previous reporting. Their standards of evidence are different too. But they share one increasingly scarce resource:
human attention.
Search itself has become easier. Databases are larger. Alerts are ubiquitous. AI can retrieve, rank and summarise information at a scale that would have been impractical only a few years ago. That does not necessarily make the underlying research problem easier; if anything, it makes the next problem more important: deciding what actually deserves attention, what can be trusted, what needs verification, what should be retained and what should still be available months later when the question changes.
I ran into this problem while trying to keep a fast-moving body of technical research useful for later analytical work. I did not need a formal systematic review every few weeks, and I did not need an autonomous research system that claimed to decide what was true. I needed something more modest: a way to remain reasonably current, preserve material that might matter, understand where it came from and make the accumulated research useful when I returned to a specific problem.
What became useful was a maintained research corpus: not a perfect record of the field, not a truth machine, and not a substitute for a systematic review, but a working body of evidence that is current enough, structured enough and traceable enough to support the actual work I want to do.
The corpus is not the task
That distinction matters. A consultant does not maintain research because maintaining research is the objective; the objective may be advising a client whether to enter a market. An analyst may be trying to understand whether a company's competitive position has changed. A journalist may be investigating whether a public claim survives contact with the underlying documents. A policy specialist may need to know whether a regulatory position is shifting. In each case, the corpus sits upstream of the real task, and its value comes from making that later work better informed.
That changes how I think about completeness. For some kinds of work, completeness is itself a major methodological requirement: a systematic review, clinical guideline, regulatory submission or other high-assurance evidence product may reasonably demand formal search strategies, explicit eligibility criteria, independent review and careful risk-of-bias assessment. A maintained corpus for continuing professional research makes a narrower claim. It is meant to help someone stay sufficiently informed for a defined purpose, and that does not make accuracy or provenance optional, but it does mean the right balance between coverage, speed and researcher effort can be different. Someone with a different purpose should make different choices.
What counts as useful evidence changes with the job
One reason I am reluctant to treat this as a universal research methodology is that the source landscape can change completely from one use case to another. Consider a consultant evaluating a possible data-centre investment. Relevant material might include:
- grid and infrastructure plans;
- regulation and tax policy;
- planning records;
- corporate disclosures;
- specialist research;
- relevant academic work.
Now consider a journalist investigating a company's claims instead. The useful hierarchy looks very different, since regulatory filings, original documents, court records, interviews, public statements, archived webpages and corroborated reporting may matter far more than academic papers. The structure of the research problem survives even though the sources do not. You still need to know:
What am I looking for?
What has changed?
What is this source?
How credible is it for this particular claim?
Does it duplicate or supersede something I already have?
Is it important enough to deserve attention now?
What should I retain?
What will I need to revisit later?
That is why I think of the maintained corpus less as a research product than as a piece of working infrastructure. The implementation may vary enormously; the attention problem does not.
More information does not automatically produce more knowledge
One response to information overload is simply to collect more, and that is easy now: alerts can accumulate indefinitely, search engines can return thousands of records, AI systems can produce long summaries of large document sets, and reference managers can store everything. But collection and knowledge are not the same thing. A corpus that grows without discipline can become another form of overload. You eventually have hundreds or thousands of items, many of them overlapping, outdated, superseded or only marginally relevant, and the effort required to understand the corpus begins competing with the work the corpus was meant to support.
That leads to a different objective:
The goal is not maximum accumulation. It is to allocate limited attention to the evidence most likely to matter, while preserving enough provenance to revisit what was deprioritised.
The aim is to retain enough of the evidence landscape that later work can be informed by it, while avoiding the fiction that every potentially relevant source deserves equal attention. That is where prioritisation becomes useful.
Prioritisation is not the same as completeness
Machine-assisted screening and active learning have become increasingly useful because they can move likely relevant material towards the front of a review queue, though that distinction matters more than it first appears: a system can be very good at helping you find relevant material early without being able to tell you safely that nothing important remains.
A 2026 study of active-learning stopping rules in ASReview illustrates the problem. Across five datasets and 35,000 simulated screenings, the amount of material that had to be screened before finding the final relevant record varied enormously, from an average of 2.9% to 76.9% depending on the dataset, and none of the tested stopping approaches consistently guaranteed that every relevant study had been found. The practical lesson is simple:
ranking can work even when stopping remains uncertain.
For a systematic review, that residual uncertainty may be unacceptable, but for a maintained professional corpus, prioritisation can still be extremely valuable. The objective may not be to certify that nothing important exists outside the corpus; it may be to make the most consequential new material visible early enough that a human can decide whether it changes the work, a different promise, and for many professional uses, a more honest one.
Large language models are useful when the task is bounded
LLMs make this more interesting because they can work on the content itself: they can screen abstracts, classify documents, identify passages, extract structured information and summarise material.
Recent evidence suggests some of these tasks are becoming genuinely useful. A 2026 systematic review found promising performance for screening, particularly with newer models, but much greater variability as tasks became more interpretive. A separate review of how LLMs are being evaluated found that reproducibility-critical safeguards against problems such as leakage and overfitting were often difficult to verify. That should not be surprising: extracting a stated sample size from a paper is one kind of task, but deciding whether the study design makes the conclusion credible is another. Finding the date a policy came into force is different from deciding whether it meaningfully changes a client's exposure, and identifying what a company said is different from deciding whether the available evidence supports what it implied.
The useful question is therefore not:
Can AI do research?
It is:
Which parts of this research process are bounded enough that automation can reduce routine work without silently becoming the authority?
That question travels well across domains. The answer will not be identical for a journalist, consultant, academic researcher or compliance analyst, but the underlying principle can remain the same.
The scarce thing is not human presence. It is human attention.
The obvious response is to say that AI can help, provided a human stays in the loop. I have become increasingly sceptical of that phrase, because it can describe almost anything: if a machine classifies 500 items and a human is required to click “approve” 500 times, there is technically a human in the loop, but that does not mean the resulting process is good. Humans make mistakes, reviewers disagree, attention declines with repetitive work, and people can become overly trusting of plausible automated recommendations. There is research showing that human reviewers make screening errors: one analysis of 25 systematic reviews found a total error rate of roughly 11% at abstract screening, and there is also a substantial human-factors literature around automation bias and over-reliance. So the goal cannot simply be to maximise the number of human approvals. The harder question is:
Where does human attention add enough value that it should not be automated away?
That question changed how I thought about the corpus. Human effort is most valuable where ambiguity, interpretation or consequence is high; routine discovery, initial prioritisation, structured extraction and other repetitive tasks may be increasingly suitable for assistance, while a surprising source, conflicting evidence, an important methodological limitation or a finding that could materially alter later analysis deserves more attention. The objective is not to remove the researcher. It is to stop spending the researcher's scarce attention as though every item were equally important.
A maintained corpus is a compromise
For my purposes, a maintained corpus became a useful compromise. At a high level, the pattern looks something like this:
Define the question → discover broadly enough → establish source and status → prioritise and inspect → preserve provenance → apply judgment and retain → refresh as the evidence changes.
None of those steps is particularly novel. What matters is how they are balanced around the use case: a different profession may use different sources, a different decision may require much stronger verification, a higher-consequence use may justify substantially more human review, and a rapidly changing field may need frequent refreshes while a slower domain may not. The general pattern remains adaptable because the authority of the evidence is contextual: a peer-reviewed meta-analysis, a company press release, a court filing and an anonymous source do not deserve the same treatment merely because they all appear in the same corpus. The corpus has to preserve those distinctions rather than flatten them.
Source types and assurance requirements change by use case. The attention problem does not.
Provenance matters because sources change
Keeping the original source attached to a finding sounds obvious until the source itself starts changing. Academic publishing provides a good example: a preprint may later become a journal article, findings may be revised, corrections may appear, and in rare cases, papers are retracted. Research methodology has learned to take these relationships seriously; Crossref, for example, explicitly supports links between preprints, accepted manuscripts, versions of record, updates, corrections and retractions.
The same general problem exists outside academia. A regulation can be amended, a government consultation can become final guidance, a planning proposal can be approved, rejected or revised, a company's preliminary result can later be replaced by an audited filing, a public statement can be corrected, and a journalist can discover that two apparently separate claims trace back to the same original source. This is why provenance is not merely clerical metadata around the research: it affects the meaning of the evidence.
A useful corpus should make it possible to ask not only:
What do I know?
but also:
Why do I think I know it, which source did it come from, what status did that source have at the time, and has anything happened since that should change my view?
That is a much higher bar than saving a URL, but it is also a much more useful foundation for later analytical work.
The trade-off does not disappear
There is an uncomfortable part to all of this. No workflow makes completeness, speed, accuracy and researcher burden optimal at the same time. Exhaustive human review consumes enormous effort and does not eliminate human error. Aggressive automation reduces burden but raises the possibility of silent misses. Prioritisation helps surface likely relevant information but does not guarantee that nothing important remains unseen. LLMs can accelerate screening and extraction but become less dependable as judgment becomes more interpretive. Sampling can help detect whether a process has gone wrong, but checking a few examples does not prove that everything unchecked is correct, and retaining everything “just in case” eventually creates its own research debt.
These are not necessarily problems that disappear when the next model arrives. Some are features of the underlying work: research means making decisions while knowledge is incomplete, evidence changes and attention is finite, and any system supporting that work inherits those constraints.
The assurance level should follow the use
This is why I would not prescribe one level of human review or one corpus design for everybody. A medical systematic review should have different assurance requirements from an analyst tracking a technology market. A journalist making a serious allegation may require stronger corroboration than when monitoring a developing story, and a consultant preparing high-level market intelligence may tolerate uncertainty that would be unacceptable in technical due diligence supporting a billion-dollar investment.
The corpus should therefore be understood in relation to what comes after it: the more consequential the downstream decision, and the more heavily it relies on the corpus, the stronger the research controls should become. For my own use, the corpus is deliberately upstream. It helps determine what I should examine and what existing research may matter to a particular question, but it does not make the final decision for me. That boundary is important.
From information overload to attention allocation
I started with what looked like an information-overload problem. I now think that description misses the more useful point: access to information is already extraordinary, and the harder problem is deciding where serious attention should go.
A maintained corpus can help turn an undifferentiated stream into something more workable. It can help surface evidence likely to matter, keep a preprint distinguishable from a later version, and preserve where a conclusion came from. It can retain uncertainty instead of turning every extracted statement into a fact, and let repetitive machine assistance remain assistance rather than silently becoming judgment. And, perhaps most importantly, it can stop scarce human attention being spent uniformly across work whose value is anything but uniform.
That does not solve information overload; it does something more modest. For the kinds of work I care about, it makes a changing evidence landscape manageable enough that the actual analysis can remain informed by research. Sometimes that is all the infrastructure needs to do.