The hardest part of a literature review isn't the reading volume-it's that you forget the first paper by the time you reach the tenth, and the review you finally write reads like a patchwork of abstracts with no thread, no comparison, no gap analysis. Telling AI to "write me a literature review" is worse: it fabricates references that look disturbingly real-author names, journal titles, years, even DOIs-and your submission gets bounced. The problem is always the instruction: no reading goal, no output structure, no guardrail against fabricated citations. This pack goes from reading a single paper to comparing many to drafting a full review, in three levels, plus a 5-minute skim cheatsheet. The key is giving the model "role + task + constraints + output format" at every level, and hammering the citation-fabrication issue repeatedly.
Beginner: Structured Reading of a Single Paper
You've just gotten a paper or long article. Don't aim for exhaustive coverage yet-nail the six elements: research question, method, key data, conclusion, limitations, and a one-line summary. This is the atomic unit of any review; if a single paper is misread, every downstream comparison is garbage in, garbage out.
You are an academic reading assistant. Read the following paper/article and produce structured reading notes.
Requirements:
1. Research question: What is the author trying to answer? One sentence
2. Method: What method/dataset/experimental design was used? 2-3 sentences
3. Key data: List 3-5 core numbers or findings, keeping original units
4. Conclusion: What is the author's main conclusion? 2-3 sentences
5. Limitations: Limitations the author acknowledges + potential ones you identify
6. One-line summary: Summarize the full text in ≤20 words for future retrieval
Constraints:
- Do not include anything not mentioned in the source
- Keep original wording for key terms; gloss in parentheses if needed
- If the source lacks data, explicitly mark "not provided in source"-do not fill gaps
Input:
- Paper title: {{paper_title}}
- Abstract/full text: {{abstract_or_full_text}}
Output format: Markdown, numbered 1-6 as aboveThe goal here isn't "summarization"-it's structured extraction. With a fixed output format, you can batch-process dozens of papers and have uniform material to feed into the comparison stage.
Intermediate: Multi-Paper Comparison and Synthesis
Reading one paper is just the start; the value of a review lies in comparison. Feed multiple papers and have AI produce a comparison table plus consensus/divergence analysis. This level is where AI's "hedging" tendency shows most: it defaults to "the studies each have their own focus." You have to force it to make disagreements concrete.
You are a literature comparison analyst. Read the following papers and produce a comparison table and synthesis.
Input papers:
1. {{Paper 1: title/author/year}}
2. {{Paper 2: title/author/year}}
3. {{Paper 3: title/author/year}}
(add more as needed)
Task 1: Comparison table
Markdown table with these dimensions:
| Paper | Research question | Method | Data/sample | Core conclusion | Limitations |
One row per paper, trimmed for table readability.
Task 2: Consensus and divergence
- Consensus (≥2 papers agree): List 3-5 points, noting which papers agree
- Divergence: List 2-3 points, stating "Paper A finds X, Paper B finds Y"-do not paper over with "each has its focus"
- Likely cause of divergence: Data difference? Method difference? Definition difference? Give your judgment
Task 3: Synthesis insight
- What conclusion do these papers collectively point toward?
- What questions do none of them answer? (Initial gap identification-deeper analysis at expert level)
Constraints:
- Analyze only the provided papers; do not introduce external literature
- If two papers reach the same conclusion with different wording, classify as "consensus" and note the wording difference
- Table entries must be traceable to the source; do not distort through over-summarizationThis output feeds directly into the "thematic comparison" section of a review. The key is fixing the table dimensions-so the framework works across topics and becomes a reusable analytical template.
Expert: Full Literature Review Structure
At this level, you're not summarizing what exists-you're scaffolding a review. Given a topic, have AI produce a complete review structure, identify research gaps, and suggest citation directions. This is the most dangerous level for fabricated citations, so you need three layers of guardrails.
You are a literature review methodology expert. Based on the given topic, produce a review structure draft.
Topic: {{research_topic}}
Existing literature list (optional): {{literature_list, or let AI draft a framework from general knowledge}}
Review goal: {{e.g., prepare a thesis proposal / write a journal review / internal tech survey}}
Task 1: Review structure
Lay out the following skeleton, with 2-3 sentences per section describing what it should cover:
1. Background: Why does this field matter? Define core concepts
2. Research thread: The evolution of the topic, organized chronologically or by method school
3. Sub-themes (3-5): Split by problem dimension; list key literature directions under each
4. Research gaps: What haven't existing studies answered? Classify as "data gap / method gap / theory gap"
5. Trends: Likely directions in the next 1-3 years
Task 2: Gap identification
- List 3-5 gaps, each with: gap description + why it's a gap (which papers come close but don't cover it) + potential research value
- Prioritize: which gap is most worth pursuing
Task 3: Citation suggestions
- For each sub-theme, list 3-5 "suggested citation directions" (e.g., "empirical studies on X post-2023")
- Clearly distinguish: which are directions AI is fairly confident exist in its knowledge vs. which are speculative
- Label: The following are directional suggestions based on AI knowledge. All citations must be verified against original sources before use.
Hard constraints:
- Do not fabricate specific references (title/author/journal/year). For examples, use directional descriptions, e.g., "search NeurIPS 2023 papers on X"
- This draft is scaffolding, not a finished product. Mark sections that need human follow-up
- All citations must be verified against original sources; AI-provided directions serve only as search leads
Output: Markdown, sectioned by Task 1/2/3The mindset at expert level: AI builds the scaffolding, you lay the bricks. It's good at structuring, gap-finding, and threading the narrative, but it cannot guarantee that a cited paper actually exists. Write the "no fabrication" rule into the prompt's hard constraints, then verify every reference against a database afterward.
Cheatsheet and Recommendations
5-Minute Skim Cheatsheet
In a hurry and just need to decide whether a paper is worth a full read? Use this minimal prompt:
You are a speed-reading assistant. Extract the core of a paper in 5 minutes using this structure:
1. One sentence: What problem does this paper solve?
2. One sentence: What method does it use?
3. One sentence: What is the most important finding/number?
4. One sentence: What is the conclusion?
5. Reader decision: Worth a full read / Skim is enough / Skip. Give one reason
Input: {{paper title + abstract}}
Output: 5 lines, one answer per line, ≤15 words/lineGeneral Constraints (Apply to Every Level)
- No fabricated citations-This is the #1 risk of AI-assisted reviews. AI generates references that look alarmingly real: plausible author names, real journal titles, correct-looking years, even fake DOIs. Write "all citations must be verified against original sources" into every prompt, then check each one in Google Scholar or a database. When in doubt, delete.
- Distinguish AI output from human review-An AI-produced review is a draft, not a final product. It can scaffold structure, find gaps, and run comparisons, but academic judgment, citation accuracy, and logical soundness ultimately rest with you. Label the output "AI-generated draft, requires human verification"-don't submit AI scaffolding as final work.
- De-AI the prose-Common AI review tells: overuse of "it is worth noting that" and "in summary"; every paragraph starts with "firstly/secondly/finally"; comparisons default to "each study has its own focus" to dodge real disagreement. Add "avoid mechanical transition phrases; don't start every paragraph with firstly/secondly/finally; make disagreements concrete" to the prompt, or revise manually after output.
Recommended Model Stack
- Long-context reasoning: Claude (Opus/Sonnet) or DeepSeek-large context window, stable logic chains, well-suited to expert-level review scaffolding and single-paper deep reads.
- Web-connected retrieval: Kimi or Perplexity-have AI search for real literature first, then analyze, to cut fabrication risk. Use Perplexity for literature discovery, then Claude/DeepSeek for deep reasoning.
- Multi-paper comparison: Prioritize models with large context windows (Claude 200K, DeepSeek, Kimi long-text); comparing 3-5 full papers at the intermediate level requires the room.
References
- Anthropic Prompt Engineering Guide-the official methodology for structured prompts; the role/task/constraints/output format paradigm
- OpenAI Prompt Engineering Guide-six-element strategies and prompt best practices
- PRISMA Statement-the international standard for systematic review reporting; use as a checklist when verifying review structure
- Cochrane Handbook for Systematic Reviews-the authoritative manual for systematic review methodology; the gap identification and evidence grading chapters are especially useful