agent-pages

Research Program Plan for an Unspecified Topic

Download

Research Program Plan for an Unspecified Topic

Executive summary

Because no subject has been fixed, the most defensible approach is to treat topic selection as the first research phase rather than quietly choosing a theme and building an elaborate review around an untested assumption. The proposed program begins with a scored portfolio of candidate topics, selects one using explicit criteria, converts it into a bounded research question, and then scales the work to the available horizon: a one-week evidence scan, a four-week rapid review with reproducible analysis, or a twelve-week mixed-method study.

The default quality standard is primary and official evidence first. Peer-reviewed experiments, administrative data, legislation, regulatory filings, statistical-agency datasets, and original technical documentation should precede consultancy reports, news coverage, vendor material, and unsourced summaries. The uploaded research brief's emphasis on distinguishing measured outcomes from inference, recording study design and limitations, and labeling evidence strength is adopted here as a reusable standard.

The three highest-priority candidate topics are:

  1. The effects of generative AI on knowledge-work productivity, quality, and skill development.
  2. The unequal health and economic consequences of urban heat exposure.
  3. The effectiveness and equity of AI-supported tutoring and personalized learning.

These topics rank highly because they combine societal impact, stakeholder interest, cross-disciplinary relevance, and the possibility of triangulating published research with official or administrative data. They are also broad enough to support different project scales but can be narrowed to a particular occupation, city, country, age group, intervention, or outcome.

A rapid review should still be systematic. The search strategy, inclusion rules, screening decisions, and analytical shortcuts should be recorded before synthesis begins; PRISMA 2020 and PRISMA-S provide useful reporting structures, while Cochrane's rapid-review guidance emphasizes making time-saving departures from a full systematic review explicit rather than invisible.

The operational workflow is:

flowchart LR
    A[Candidate topic portfolio] --> B[Weighted topic scoring]
    B --> C[Stakeholder and feasibility check]
    C --> D[Bounded research question]
    D --> E[Protocol and search strategy]
    E --> F[Literature and data retrieval]
    F --> G[Screening and quality appraisal]
    G --> H[Qualitative or quantitative analysis]
    H --> I[Triangulation and uncertainty assessment]
    I --> J[Report, evidence tables, data products]

Research portfolio and selection criteria

A strong topic is not merely interesting. It must be important enough to justify attention, narrow enough to investigate, and supported by data capable of answering the proposed question rather than merely illustrating it.

The following weighting scheme can be used for initial prioritization. Scores should be assigned independently by at least two reviewers or stakeholders when possible, followed by a short reconciliation discussion.

Criterion Suggested weight Assessment question High-score indicators
Potential impact 25% Could the findings materially affect decisions, welfare, risk, cost, or capability? Large affected population, high cost of error, actionable intervention
Feasibility 20% Can a credible answer be produced with the available time, skills, access, and ethics constraints? Bounded scope, accessible methods, manageable approvals
Data availability 20% Are reliable literature and primary or official data available at the required level of detail? Documented variables, usable identifiers, temporal and geographic coverage
Novelty and unresolved uncertainty 15% Is there a consequential gap, contradiction, or underexplored context? Conflicting studies, new technology, weak external validity
Stakeholder interest 15% Is there a decision-maker or affected group likely to use the result? Named users, impending decision, policy or operational relevance
Ethical and legal manageability 5% Can the work be completed without disproportionate privacy, safety, or consent risks? Public or de-identified data, low-risk recruitment, clear governance

A topic should normally be rejected or redesigned when it scores poorly on feasibility or data availability even if its social importance is high. A beautiful question with no credible identification strategy tends to produce an opinion essay rather than research.

Prioritized candidate topics

The scores below are provisional planning judgments rather than empirical findings. They assume English-language research, access to common scholarly databases, and no proprietary organizational dataset.

Priority Concrete topic Domains Provisional score Rationale
1 How generative AI changes productivity, work quality, and skill development among knowledge workers Technology, economics, business, social science 91 This topic supports comparisons across occupations, task types, experience levels, and implementation models rather than treating "AI use" as a single intervention. It also allows a valuable separation between immediate output gains and longer-term effects on expertise, judgment, error detection, and worker autonomy.
2 The distributional effects of urban heat on health, labor productivity, and household vulnerability Environment, health, economics, policy 89 Heat exposure can be studied by combining weather, land-surface, health, labor, and socioeconomic datasets at several geographic scales. The topic is especially decision-relevant when narrowed to which neighborhoods or occupations bear the greatest burden and which adaptation measures provide measurable protection.
3 Whether AI-supported tutoring improves durable learning and educational equity Education, technology, social science, policy 87 This question distinguishes learning outcomes from engagement metrics such as sessions, prompts, or time on platform. It can also test whether effects vary by prior attainment, teacher involvement, language, disability, socioeconomic status, and the quality of instructional design.
4 Which antimicrobial-stewardship interventions reduce inappropriate prescribing without harming access or outcomes Health, policy, behavioral science 84 The topic has clear intervention–outcome structures and can draw on trials, prescribing records, surveillance systems, and implementation studies. Its main analytical challenge is separating genuine improvement from diagnostic substitution, under-recording, or reduced access to needed treatment.
5 How renewable-energy penetration affects grid reliability, storage needs, and consumer costs Environment, technology, economics, policy 82 The question can be narrowed by market, generation mix, weather exposure, storage configuration, or regulatory design. It supports scenario analysis and observational comparisons, although causal claims require care because infrastructure investment and policy adoption are not randomly assigned.
6 How housing affordability affects labor mobility, family formation, and access to high-productivity regions Economics, policy, social science 81 Official housing, wage, migration, and household data can support longitudinal and cross-regional analysis. The work becomes most actionable when it compares concrete policy mechanisms such as zoning reform, social housing, transport investment, rent regulation, or housing subsidies.
7 Whether digitalization improves productivity and resilience among small and medium-sized enterprises Business, technology, economics 79 This topic can examine specific capabilities such as cloud adoption, digital payments, e-commerce, cybersecurity, or enterprise software rather than using a vague digital-maturity score. It has strong stakeholder relevance, but self-selection is a major concern because better-managed firms may both digitalize earlier and perform better.
8 The relationship between social-media use patterns and adolescent mental health Health, education, social science, policy 77 The research should distinguish active communication, passive consumption, social comparison, harassment, sleep displacement, and compulsive use instead of relying only on total screen time. Strong conclusions will require longitudinal, experimental, or quasi-experimental evidence because cross-sectional associations are highly vulnerable to reverse causality and confounding.
9 Cybersecurity maturity and incident resilience in critical infrastructure or public-sector organizations Technology, policy, business 75 The topic can connect governance practices, staffing, asset management, incident response, and supplier risk to observable resilience indicators. Data access and disclosure bias are substantial obstacles because organizations with severe weaknesses or incidents may be least willing to participate.
10 Whether climate-adaptation finance reaches vulnerable populations and produces measurable resilience Environment, economics, policy, social science 74 This topic can trace funds from commitments through disbursement, implementation, and local outcomes, revealing where nominal spending does not translate into protection. Attribution is difficult because adaptation projects are heterogeneous and benefits often emerge slowly or only when an extreme event occurs.

The portfolio intentionally covers all requested domains while favoring questions that can be addressed through more than one evidentiary route. A final selection should pass four gates: a definable population or system, a measurable exposure or intervention, a meaningful outcome, and a plausible comparison or counterfactual.

Rapid review and evidence architecture

The initial review should be designed to answer a bounded decision question, not to accumulate everything written about a broad theme. The topic should therefore be translated into a structure appropriate to the field:

Research type Recommended framing
Intervention or program evaluation Population, intervention, comparator, outcomes, time horizon, setting
Exposure or risk-factor research Population, exposure, comparator, outcomes, confounders
Policy analysis Jurisdiction, policy instrument, affected population, implementation mechanism, outcomes
Technology assessment Users, task, technology configuration, comparator, performance, safety and organizational outcomes
Qualitative inquiry Population, phenomenon of interest, context, stakeholder perspective
Business or market research Unit of analysis, strategic change, comparison group, operational and financial outcomes

A one-page protocol should be written before broad searching. It should state the question, definitions, databases, dates, languages, study designs, outcomes, quality criteria, planned synthesis, and any intentional rapid-review shortcuts. PRISMA provides reporting checklists and flow diagrams, while PRISMA-S focuses specifically on transparent reporting of literature searches.

Search process

The literature search should use three layers.

The discovery layer identifies terminology, seminal works, major authors, reviews, and citation clusters. OpenAlex can support broad discovery and citation-network exploration, while Crossref provides openly retrievable scholarly metadata and relationships such as publication updates, funders, licenses, and linked research objects.

The disciplinary layer searches specialized databases using subject headings and field-specific vocabulary. Examples include PubMed for biomedical and life-science literature, ERIC for education, PsycINFO for psychology, ACM Digital Library and IEEE Xplore for computing, and EconLit or RePEc for economics. PubMed is maintained by the U.S. National Library of Medicine and includes citations and abstracts rather than guaranteeing full-text access.

The verification layer checks the original article, official report, trial registration, dataset documentation, code repository, correction notices, and retraction status. Aggregated abstracts and AI-generated summaries should never be the final evidentiary source for an important claim.

A reusable query template is:

(population OR setting synonyms)
AND
(intervention OR exposure OR technology synonyms)
AND
(outcome synonyms)
AND
(study-design filters, only when justified)

For the generative-AI topic, an initial query family might combine terms for generative AI or large language models with productivity, quality, error, expertise, learning, employment, field experiment, randomized trial, or longitudinal study. For urban heat, the blocks would cover extreme heat, temperature, heatwave, mortality, morbidity, labor productivity, exposure inequality, adaptation, and spatial or panel methods.

Inclusion and exclusion rules

A practical default is to include:

Dimension Default inclusion rule
Language English-language material, while recording the likely direction of language bias
Publication period Recent ten years, plus earlier seminal studies and foundational methods
Evidence type Primary empirical studies, systematic reviews used for discovery, official statistics, legislation, standards, technical documentation, and transparent working papers
Population Directly relevant populations, with adjacent populations retained only when transferability is discussed explicitly
Outcomes Predefined behavioral, learning, health, environmental, economic, operational, or policy outcomes
Methods Studies with enough information to evaluate design, measurement, analysis, and limitations
Geography Global initially, narrowed when institutional context materially affects interpretation
Access Full text preferred; abstract-only records may remain in the search log but should not support detailed conclusions

Default exclusions should include duplicate reports of the same study, undated promotional material, commentary without original evidence, datasets lacking provenance or definitions, studies whose outcomes do not match the research question, and sources that make causal claims from purely descriptive comparisons without adequate justification.

Screening should occur in two passes: title and abstract, followed by full text. In a one-person rapid review, a second reviewer can independently screen a random sample and all borderline records; in a larger project, dual screening is preferable for the complete set. The search log should record the query, platform, date, filters, result count, deduplication process, and reasons for full-text exclusion.

Methodology selection

Intended inference Suitable quantitative methods Suitable qualitative methods Mixed-method contribution
Describe scale or distribution Descriptive statistics, rates, trends, spatial mapping Stakeholder accounts of where official measures miss lived conditions Explain discrepancies between measured prevalence and reported experience
Estimate causal effect Randomized trial, difference-in-differences, event study, regression discontinuity, instrumental variables, synthetic control Process tracing of implementation and alternative mechanisms Test effect estimates and explain why effects differ across settings
Understand mechanisms Mediation, moderation, sequence analysis, process measures Interviews, observation, thematic or framework analysis Link statistical heterogeneity to organizational or behavioral processes
Compare policies or systems Multilevel models, matched comparisons, panel analysis Comparative case study, document analysis Identify both outcome differences and institutional conditions
Forecast or plan scenarios Time series, microsimulation, system dynamics, Bayesian models Expert elicitation, scenario workshops Combine model outputs with institutional feasibility and uncertainty
Evaluate implementation Adoption, fidelity, reach, cost, and outcome models Interviews with implementers and affected groups Separate intervention failure from implementation failure

Randomized evidence should not automatically outrank all observational evidence. The design must match the question: a small, short laboratory trial may estimate a narrow treatment effect well but offer weaker external validity than a carefully designed administrative-data study covering years and diverse populations.

Evidence grading

Each conclusion should be labeled using a transparent scheme:

The review must distinguish absence of evidence from evidence of no effect. It should also separate intermediate metrics—such as adoption, clicks, attendance, or engagement—from final outcomes such as durable learning, reduced mortality, higher-quality work, or lower total cost.

Detailed plans for the leading topics

The three plans below are scoped as four-week rapid-research projects. Each can later be extended into a twelve-week causal or mixed-method study.

Plan element Generative AI and knowledge work Urban heat, health, and productivity AI-supported tutoring and learning
Core objective Estimate how generative-AI assistance affects completion time, output quality, error rates, skill development, and worker experience, and identify which task and worker characteristics moderate those effects. Quantify how heat exposure affects health and labor outcomes across neighborhoods or occupational groups, and assess the evidence for specific adaptation measures. Determine whether AI tutoring improves durable learning, transfer, and equity compared with conventional instruction or non-adaptive digital tools.
Primary scope Knowledge workers performing writing, analysis, coding, customer support, or related cognitive tasks; the scope should ultimately be narrowed to one or two occupations. One country or a set of comparable cities, with a defined heat measure, health or labor outcome, and vulnerability framework. A bounded educational stage and subject, such as secondary mathematics, introductory programming, or language learning.
Main questions How large are short-run productivity and quality effects? Do effects differ by prior expertise, task complexity, verification requirements, or degree of workflow integration? What is known about longer-term skill formation, deskilling, autonomy, and error propagation? What is the exposure–response relationship? Which populations experience the greatest burden? Do cooling centers, tree cover, building retrofits, work-rest rules, warnings, or household cooling access reduce that burden? Do systems improve post-test and delayed-test performance? Are gains concentrated among initially lower- or higher-performing learners? How do teacher oversight, feedback quality, hallucinations, and curriculum alignment affect outcomes?
Working hypotheses AI will produce larger immediate time gains on bounded tasks than on tasks requiring tacit judgment, contextual knowledge, or independent verification. Less-experienced workers may obtain larger short-run assistance benefits but may also face greater difficulty detecting subtle errors. Health and productivity burdens will rise nonlinearly at high temperatures and will be larger where housing, occupational, health, and infrastructure vulnerabilities overlap. Adaptation effects will vary with access, actual usage, maintenance, and behavioral response rather than infrastructure availability alone. Systems aligned with curriculum and supported by teachers will produce more reliable learning gains than unconstrained conversational tools. Engagement improvements will not consistently translate into delayed retention or transfer unless practice, feedback, and difficulty progression are instructionally sound.
Priority literature Randomized and field experiments, longitudinal workplace studies, human–AI interaction research, software-engineering studies, labor-economics research, implementation studies, and official workforce statistics. Epidemiological studies, environmental-health research, labor-economics studies, urban-climate analysis, adaptation evaluations, and official meteorological, health, census, and labor data. Randomized and quasi-experimental educational evaluations, learning-science studies, intelligent-tutoring-system research, teacher-implementation studies, and official assessment and participation data.
Primary or official data sources National labor-force and occupational datasets, official productivity statistics, occupational taxonomies, organizational workflow logs where access exists, and original experiment data or supplements. National meteorological services, Copernicus or comparable reanalysis data, satellite-derived land-surface measures, census data, health records, mortality statistics, labor-force surveys, and city infrastructure data. UNESCO UIS, national education ministries, assessment agencies, school administrative records, classroom or platform logs, and original trial repositories. UNESCO UIS supplies internationally comparable education, science, and cultural statistics and maintains a broad education data browser.
Review method Rapid systematic review with separate evidence maps for productivity, quality, skill, and worker outcomes. Meta-analysis should be attempted only when interventions, tasks, comparators, and outcomes are sufficiently comparable. Rapid evidence review plus a documented data-availability assessment. If suitable data can be assembled quickly, add a small spatial or panel-analysis demonstration rather than presenting it as definitive causal evidence. Rapid systematic review with separate synthesis for adaptive tutors, generative conversational systems, teacher-facing tools, and automated feedback. Separate immediate post-tests, delayed retention, transfer, engagement, and equity outcomes.
Quantitative analysis Extract standardized mean differences, percentage changes in completion time, error or defect rates, and subgroup interactions. Use random-effects synthesis where defensible, prediction intervals when enough studies exist, and sensitivity analysis excluding vendor-authored or non-peer-reviewed work. Construct exposure and outcome definitions; calculate descriptive gradients by vulnerability group; consider distributed-lag or panel models if temporal data are adequate. Use spatial cross-validation and sensitivity tests for temperature metric, lag structure, seasonality, and geographic aggregation. Extract standardized learning effects, baseline-adjusted outcomes, attrition, delayed-test effects, and subgroup differences. Conduct moderator analysis for subject, age, duration, teacher involvement, adaptivity, assessment type, and risk of bias when sample size permits.
Qualitative analysis Code reported mechanisms: task decomposition, verification, trust, autonomy, workflow redesign, surveillance, training, and managerial incentives. Analyze policy and implementation documents for target population, eligibility, reach, maintenance, behavioral assumptions, and documented barriers. Code teacher and learner evidence for trust, feedback usefulness, cognitive load, accessibility, cheating concerns, curriculum fit, and implementation burden.
Triangulation Compare controlled experiments with real-world implementation studies; treat disagreement as evidence about context rather than averaging unlike settings. Compare modeled exposure, administrative outcomes, and stakeholder-reported vulnerability; document where official measures omit informal work or housing conditions. Compare platform engagement, formal assessment results, and teacher or learner accounts; do not infer learning from usage alone.
Expected deliverables A twenty-page analytical report, study and effect-size table, task-by-worker evidence matrix, implementation-risk register, reproducible extraction file, and a decision framework for responsible deployment. A twenty-page report, exposure and vulnerability map prototype, policy-intervention evidence table, reproducible data dictionary, and prioritized adaptation options with uncertainty labels. A twenty-page report, intervention taxonomy, learning-outcome evidence table, equity and safety matrix, procurement questions for institutions, and reproducible extraction file.
Main risks Fast-moving technology definitions, publication lag, vendor incentives, heterogeneous tasks, weak evidence on long-term skill effects, and selective reporting of successful deployments. Exposure misclassification, ecological fallacy, geographic data mismatch, undocumented adaptation behavior, missing informal-sector data, and causal confounding. Short interventions, non-comparable tests, attrition, novelty effects, privacy constraints, vendor-reported engagement metrics, and sparse delayed-retention evidence.
Week one milestone Finalize occupations and outcomes; preregister protocol; search and screen reviews, experiments, and official workforce sources; construct taxonomy of tasks and AI configurations. Select geography and outcomes; define heat metrics and vulnerability variables; retrieve foundational studies and official datasets; assess spatial and temporal compatibility. Select educational stage and subject; finalize intervention taxonomy and outcomes; retrieve trials, reviews, and official education data; define equity subgroups.
Week two milestone Complete primary-study screening and extraction; conduct quality appraisal; retrieve available supplements or datasets; begin quantitative harmonization. Complete evidence screening; clean a small demonstrator dataset; create initial exposure and vulnerability summaries; appraise adaptation studies. Complete study screening and extraction; appraise bias and measurement quality; code assessment timing, instructional design, and implementation characteristics.
Week three milestone Analyze effects and heterogeneity; synthesize mechanism evidence; compare experimental and field evidence; draft findings and caveats. Run descriptive, spatial, or panel analyses; conduct sensitivity checks; synthesize adaptation evidence and implementation constraints. Calculate comparable learning effects where possible; analyze moderators; synthesize teacher, learner, safety, and implementation evidence.
Week four milestone Complete report, evidence tables, decision matrix, reproducibility package, and stakeholder presentation; perform source and claim audit. Finalize maps, findings, adaptation matrix, uncertainty discussion, technical appendix, and presentation; audit geographic and causal claims. Complete report, intervention taxonomy, evidence and equity tables, procurement guidance, reproducibility package, and claim audit.

The top three plans share an important boundary: a four-week project can produce a credible evidence synthesis and exploratory secondary analysis, but it should not promise definitive answers to long-term causal questions when the existing data do not support them.

Timelines, resources, deliverables, and risk controls

The project scale should be chosen according to the decision at stake. A one-week scan is appropriate for prioritization or an early meeting; a four-week project is suitable for a serious rapid review; a twelve-week project can add deeper synthesis, original analysis, stakeholder research, and stronger validation.

Project horizon Appropriate purpose Estimated effort Indicative non-salary budget Core activities Deliverables
One week Topic selection, executive briefing, data feasibility, initial decision support 30–60 person-hours USD 0–2,500 Scope definition, targeted searches, 25–50 high-value sources, official-data inventory, preliminary evidence map, key expert or stakeholder consultation if readily available Five-to-ten-page scoping brief, candidate-question matrix, search log, source library, risk register, recommendation on whether to proceed
Four weeks Rapid evidence review with reproducible descriptive or exploratory analysis 140–300 person-hours USD 2,500–20,000 Protocol, systematic searches, screening, structured extraction, quality appraisal, quantitative or qualitative synthesis, limited secondary-data analysis, internal review Fifteen-to-thirty-page report, evidence tables, data dictionary, analytical notebook, charts, stakeholder deck, limitations and research-gap map
Twelve weeks Decision-grade mixed-method study, stronger causal analysis, or publication-oriented review 600–1,200 person-hours USD 20,000–150,000 or more Expanded review, preregistration, larger data engineering effort, interviews or surveys, advanced causal or spatial analysis, robustness testing, external expert review Full technical report, executive brief, cleaned dataset where licensable, reproducible code, dashboard or interactive visualizations, policy or implementation playbook, manuscript draft

These budget ranges are planning estimates, not market quotations. Data licensing, participant recruitment, transcription, translation, cloud computing, secure environments, ethics review, specialist consulting, and geographic fieldwork can move a project far outside the indicated ranges.

A typical four-week program can be visualized as follows:

gantt
    title Four-week rapid research program
    dateFormat  YYYY-MM-DD
    axisFormat  %b %d

    section Framing
    Topic scoring and scope            :a1, 2026-08-03, 2d
    Protocol and definitions           :a2, after a1, 2d

    section Retrieval
    Database and official-source search :b1, 2026-08-05, 6d
    Deduplication and first-pass screen :b2, 2026-08-07, 6d
    Full-text screening                 :b3, after b2, 5d

    section Evidence
    Extraction and quality appraisal   :c1, 2026-08-12, 9d
    Data acquisition and cleaning      :c2, 2026-08-10, 10d

    section Analysis
    Quantitative or qualitative analysis :d1, 2026-08-18, 7d
    Triangulation and sensitivity checks :d2, after d1, 3d

    section Delivery
    Draft report and visualizations    :e1, 2026-08-23, 5d
    Claim audit and final package      :e2, after e1, 3d

Resource model

A lean four-week team would contain one lead researcher, one analyst or research assistant, and limited subject-matter review. A stronger team would add an information specialist for search design, a statistician or methodologist, a domain expert, and a stakeholder or implementation researcher.

Recommended tools include:

Function Lean option Expanded option
Reference management Zotero Zotero group library with controlled tagging and review roles
Search and discovery Subject databases, OpenAlex, Crossref, citation chasing Licensed Scopus or Web of Science access plus librarian-designed searches
Screening Spreadsheet or Zotero tags Rayyan, Covidence, EPPI-Reviewer, or equivalent
Extraction Structured CSV or spreadsheet REDCap, Airtable, DistillerSR, or a validated review platform
Quantitative analysis R or Python R, Python, Stata, or specialized meta-analysis packages
Qualitative analysis Structured coding matrix NVivo, MAXQDA, ATLAS.ti, or reproducible text-analysis workflows
Reproducibility Git and documented folders Git repository, OSF registration, data versioning, containers
Visualization R, Python, spreadsheet charts GIS, interactive dashboard, network-analysis tools
Project management Markdown checklist and calendar Issue tracker, project board, decision log, responsibility matrix

Risk and mitigation register

Risk Consequence Mitigation
Scope expansion Shallow treatment of many questions and missed deadlines Freeze the primary question and outcomes in the protocol; place adjacent questions in a research backlog
Weak or inaccessible data Descriptive claims cannot be tested Complete a data-feasibility gate in the first week; prepare an alternative review-only design
Heterogeneous definitions Invalid pooling or misleading comparison Build a concept dictionary and intervention taxonomy before extraction
Confounding and reverse causality Association is misreported as effect Draw a causal diagram, identify the counterfactual, and label descriptive versus causal analyses explicitly
Publication and reporting bias Effects appear larger or more consistent than they are Search trial registries, preprints, working papers, dissertations, and official evaluations; inspect selective outcome reporting
Vendor or sponsor influence Engagement or favorable outcomes are overemphasized Record funding and conflicts; analyze independent and vendor-authored studies separately
Stale evidence Conclusions fail to reflect a fast-changing field Record final search date, search recent preprints and registrations, and provide an update protocol
Privacy and ethics Harm to participants or legal noncompliance Prefer de-identified data, minimize collection, document lawful basis, restrict access, and seek appropriate ethics review
English-language restriction Relevant regional evidence is missed State the restriction, search English abstracts from regional databases, and flag contexts where language bias is likely material
AI-assisted extraction errors Fabricated or distorted claims enter the report Require source-level verification, page references, structured audit fields, and human approval for every substantive claim
Stakeholder pressure Unfavorable evidence is omitted or softened Predefine outcomes, maintain a decision log, separate findings from recommendations, and disclose conflicts
Failed reproducibility Results cannot be audited or updated Preserve search exports, scripts, data dictionaries, software versions, and immutable protocol records

Literature extraction and knowledge workflow

The extraction system should be designed before full-text review. Its purpose is to prevent a common failure mode: remembering a paper's conclusion but losing the details needed to judge whether that conclusion applies.

Literature-review extraction template

Field What to record Standardized options or format Why it matters
Record ID Unique internal identifier AUTHOR_YEAR_SHORTTITLE Connects notes, PDFs, extraction rows, and analysis
Full citation Authors, year, title, venue Imported and manually verified Prevents citation errors
DOI or persistent identifier DOI, trial ID, report number, dataset ID Verified against publisher, Crossref, registry, or repository Supports deduplication and traceability
Source status Publication type Peer-reviewed; preprint; official report; working paper; vendor report; other gray literature Prevents unlike evidence from being silently treated as equivalent
Research question Authors' stated question One or two sentences Clarifies what the study actually tested
Study design Design and allocation method RCT; field experiment; quasi-experimental; cohort; cross-sectional; case study; qualitative; simulation; review Determines what inferences are supportable
Population and setting Sample size, demographics, geography, institution, occupation Structured text Determines external validity
Dates and duration Recruitment, intervention, follow-up, data period ISO dates where available Reveals short horizons and historical context
Intervention or exposure Precise definition and implementation Structured description Prevents category errors
Comparator Counterfactual or comparison condition Treatment as usual; alternative tool; no exposure; matched comparison; none Central to effect interpretation
Outcomes Definition, instrument, timing, unit Primary and secondary separated Distinguishes engagement from substantive outcomes
Main result Direction and magnitude Raw estimate and standardized estimate Provides the synthesis input
Uncertainty Confidence or credible interval, standard error, p-value where relevant Numeric Avoids relying on point estimates alone
Subgroup or moderator results Prior ability, age, sex, occupation, context, exposure level Prespecified versus exploratory Identifies heterogeneous effects
Attrition and missingness Rates, differential attrition, handling method Numeric and narrative Reveals bias risk
Analytical method Model, covariates, clustering, weights, corrections Structured text Supports methodological appraisal
Quality or bias assessment Domain-level concerns Low; some concerns; high; not assessable Feeds evidence grading
Funding and conflicts Sponsor, author affiliation, declared interests Structured text Important for interpretation
Data and code availability Repository and access conditions Open; restricted; unavailable; unclear Supports reproduction
Limitations stated by authors Authors' own caveats Short quotation or paraphrase with page Guards against overgeneralization
Reviewer limitations Additional concerns identified Structured note Captures independent appraisal
Applicability Directness to selected question Direct; partly indirect; highly indirect Prevents hidden method transfer
Key evidence note Claim that may be used in synthesis Claim plus precise page, table, or figure reference Makes final claim auditing possible
Decision Include in which synthesis Primary; contextual; background; exclude Keeps the evidence base organized

Citation and note-taking workflow

  1. Use Zotero as the bibliographic source of truth. Import records through database exports or DOI lookup, attach the full text, and correct titles, dates, author names, and publication types against the original source.

  2. Assign a stable citation key. A format such as FirstAuthorYearShortTitle should be used consistently in filenames, notes, extraction tables, code, and figure captions.

  3. Verify identifiers and publication status. Crossref exposes reusable scholarly metadata, including publication relationships and post-publication updates; DataCite provides metadata for datasets and other research outputs.

  4. Create one evidence note per source. The note should contain the research question, design, population, intervention or exposure, outcomes, numerical results, quality concerns, applicability, and page-level anchors.

  5. Separate source notes from synthesis notes. Source notes describe individual studies; synthesis notes compare studies by question, mechanism, outcome, population, or contradiction. This separation reduces the risk that an interpretation becomes confused with an author's original result.

  6. Store extraction data in a machine-readable form. CSV or a versioned database is preferable to prose-only notes because it supports deduplication, consistency checks, meta-analysis, and later updating.

  7. Preregister confirmatory decisions. OSF supports time-stamped, read-only registrations of research plans; the protocol can define hypotheses, outcomes, exclusions, and analysis before confirmatory data work begins.

  8. Version the analytical work. Search exports, screening decisions, code, data dictionaries, figures, and report drafts should be versioned. Sensitive data should not be placed in a public repository merely for the sake of reproducibility; code and synthetic or aggregated examples can be shared instead.

  9. Use a final claim audit. Every externally verifiable sentence in the deliverable should be checked against its cited source, with special attention to causal wording, population, time horizon, outcome definition, and whether the source is primary.

For a cross-disciplinary report, APA author–date style is a reasonable default, although the citation style should ultimately follow the commissioning organization or target journal. The underlying library should retain DOI and persistent-identifier metadata so that output styles can be changed without rebuilding references.

Authoritative sources and visualization plan

The source hierarchy should begin with original research and official records. Broad scholarly indexes are useful for discovery, but the final evidentiary record should point to the publisher, repository, regulator, statistical agency, trial registry, legislation database, or documented dataset whenever possible.

Sources and databases by domain

Domain Literature databases Primary and official sources to prioritize Use cautions
Cross-domain scholarship OpenAlex, Crossref, Google Scholar, Scopus, Web of Science, ProQuest Publisher records, institutional repositories, DataCite, OSF, trial or study registrations Broad indexes have different coverage and deduplication quality; verify final records against persistent identifiers. OpenAlex is an open catalog that aggregates records from sources including Crossref, PubMed, arXiv, and repositories.
Technology and computing ACM Digital Library, IEEE Xplore, arXiv, DBLP NIST publications and standards, standards bodies, patent offices, public code and benchmark repositories, original system documentation Preprints and benchmarks may not be peer reviewed; benchmark gains may not represent real-world reliability. NIST maintains searchable technical-series publications.
Health and medicine PubMed, Embase, Cochrane Library, CINAHL, PsycINFO WHO, national health ministries, CDC or ECDC, ClinicalTrials.gov and other trial registries, disease-surveillance systems, drug regulators WHO estimates may be harmonized for cross-country comparability and can differ from national estimates; metadata and methods should be examined.
Environment and climate Web of Science, Scopus, GeoRef, AGRICOLA, subject repositories IPCC, UNFCCC, Copernicus Climate Data Store, national meteorological services, NASA Earthdata, NOAA, environmental regulators Distinguish modeled, remotely sensed, station, and administrative measurements; spatial and temporal resolutions may not align
Economics and development EconLit, RePEc, SSRN, NBER working papers National statistical offices, World Bank, IMF, OECD, central banks, tax or customs authorities, labor ministries Revisions, breaks in series, purchasing-power conversions, model-based estimates, and projections must be labeled. World Bank WDI compiles development indicators from officially recognized sources; OECD Data Explorer provides access to OECD statistical data; IMF has migrated current access to its newer data portal and APIs.
Public policy and law PAIS Index, HeinOnline, SSRN, legislative research databases Statutes, regulations, official gazettes, court decisions, parliamentary records, audit offices, procurement databases, regulator decisions Policy announcements should not be confused with enactment, implementation, enforcement, or measured effects
Education ERIC, PsycINFO, Education Source, Scopus, Web of Science UNESCO UIS, OECD education data, education ministries, examination and assessment agencies, school administrative systems Enrollment, participation, completion, engagement, achievement, and durable learning are different outcomes. UNESCO UIS is an official source of internationally comparable education data.
Social sciences PsycINFO, Sociological Abstracts, Social Science Citation Index, ICPSR, OSF Census agencies, household panels, labor and demographic surveys, election studies, European Social Survey, General Social Survey Self-report, common-method bias, attrition, nonresponse, and changing question wording require explicit treatment
Business and management ABI/INFORM, Business Source, SSRN, EconLit SEC EDGAR, other securities regulators, audited annual reports, earnings filings, patent and trademark offices, labor and industry statistics Company presentations and adjusted metrics are interested communications, not neutral evidence. EDGAR supports full-text filing search, current submissions, company histories, and structured filing data.

Recommended visualizations

The visualization should follow the analytical question rather than decorate the report.

Research need Recommended visualization Interpretation safeguards
Show the evidence landscape Evidence-gap matrix by intervention and outcome Display study counts and quality separately
Explain study selection PRISMA-style flow diagram Record exact exclusion reasons and deduplication
Compare estimated effects Forest plot Include intervals, study weights, heterogeneity, and prediction intervals where appropriate
Show differences across groups Dot-and-whisker subgroup chart Label exploratory analyses and avoid unsupported rank ordering
Show change over time Line chart or event-study plot Mark policy dates, uncertainty, revisions, and pre-trends
Show spatial distribution Choropleth plus exposure or outcome layers Use appropriate denominators and disclose spatial resolution
Explain causal assumptions Directed acyclic graph Distinguish measured, unmeasured, mediator, and collider variables
Compare candidate topics Weighted scoring heatmap Preserve raw criterion scores as well as totals
Map literature relationships Citation or co-word network Avoid treating citation counts as research quality
Summarize policy choices Intervention–evidence–cost matrix Separate observed evidence from projected feasibility
Display implementation logic Theory-of-change diagram Identify assumptions and failure points
Communicate uncertainty Confidence bands, scenario ranges, or sensitivity tornado chart Avoid false precision and label assumptions prominently

The final package should contain at least four visual layers: a topic-selection matrix, a study-selection flow, one analytical figure tied directly to the main outcome, and an evidence-strength or uncertainty display. The durable principle is that every visualization should answer a named question and preserve the uncertainty needed to interpret it responsibly.