Research Methods
Paper I · Basic Sciences. Six study modes, from notes to quick review.
Jump to a section
Study Notes
1. VALIDITY
1.1 Definition
Validity = the degree to which a test or instrument measures what it claims to measure. A valid IQ test actually measures intelligence, not reading speed.
1.2 Types of Validity
| Type | Definition | Example in Psychiatry |
|---|---|---|
| Face | Appears to measure what it claims (subjective judgment) | PHQ-9 "looks like" a depression questionnaire |
| Content | Covers the full domain of the construct | HAM-D covers mood, guilt, insomnia, psychomotor, somatic, the full depression domain |
| Criterion | Correlates with an external gold standard | MMSE score correlates with clinical dementia diagnosis |
| Construct | Measures the theoretical construct accurately | BDI correlates with other depression measures (convergent) but not with anxiety measures (discriminant) |
1.3 Criterion Validity: Two Subtypes
| Subtype | Timing | Example |
|---|---|---|
| Concurrent | Test and criterion measured at the same time | New depression scale given alongside HAM-D to same patients |
| Predictive | Test predicts a future outcome | GRE score predicting future academic performance |
1.4 Construct Validity: Two Subtypes
| Subtype | Definition | Example |
|---|---|---|
| Convergent | Correlates with measures of the same construct | BDI correlates highly with PHQ-9 (both measure depression) |
| Discriminant | Does NOT correlate with measures of different constructs | BDI should not correlate highly with a mania scale |
1.5 Internal vs External Validity
| Feature | Internal Validity | External Validity |
|---|---|---|
| Definition | Degree to which results are due to the independent variable, not confounders | Degree to which results can be generalised to other populations/settings |
| Focus | Cause-and-effect accuracy | Applicability beyond the study |
| Stronger in | RCTs (controlled environment) | Community-based studies, pragmatic trials |
| Threats | Selection bias, maturation, testing effects, history, attrition, instrumentation, regression to mean | Sampling bias, Hawthorne effect, setting specificity, volunteer bias |
| Trade-off | Tight control increases internal but may reduce external | Broad inclusion increases external but may reduce internal |
1.6 Threats to Internal Validity (Mnemonic: SHRIMP-T)
- Selection bias, groups differ at baseline
- History, external events during study affect outcome
- Regression to the mean, extreme scores regress on retesting
- Instrumentation, measurement tool changes during study
- Maturation, natural changes over time in participants
- Practice/testing effects, repeated testing improves performance
- Testing attrition, differential dropout between groups
1.7 Threats to External Validity
- Sampling bias, non-representative sample
- Hawthorne effect, participants alter behaviour when observed
- Setting specificity, results only apply in that setting (e.g., tertiary care = community)
- Volunteer bias, volunteers differ systematically from non-volunteers
- Interaction of treatment and context, drug works in young males but not elderly females
2. RELIABILITY
2.1 Definition
Reliability = the degree to which a test produces consistent, reproducible results across time, raters, or items. A reliable thermometer gives the same reading each time you measure the same temperature.
2.2 Types of Reliability
| Type | What It Measures | Statistic | Acceptable Value |
|---|---|---|---|
| Test-retest | Stability over time (same test, same people, different times) | Pearson r or ICC | r > 0.7 |
| Inter-rater | Agreement between two or more raters | Cohen's kappa (κ) | κ > 0.6 |
| Split-half | Consistency between two halves of the same test | Spearman-Brown corrected r | r > 0.7 |
| Internal consistency | How well items within a test measure the same construct | Cronbach's alpha (α) | α > 0.7 |
2.3 Cohen's Kappa: Interpretation
Why kappa over simple percentage agreement? Kappa accounts for agreement expected by chance alone. Two raters flipping coins would agree 50% of the time, kappa corrects for this.
2.4 Cronbach's Alpha: Key Points
- Measures internal consistency, how well items "hang together"
- Range: 0 to 1 (higher = more consistent)
- If alpha < 0.7 → items are measuring different things or too few items
- If alpha > 0.95 → items may be redundant (item redundancy)
- Affected by number of items (more items → higher alpha, even if items are mediocre)
2.5 Reliability vs Validity: The Relationship
Classic analogy: A marksman shooting a tight cluster (reliable) but missing the bullseye (not valid). Scattered shots hitting the bullseye on average are neither reliable nor truly valid for individual measurement.
3. STUDY DESIGNS
3.1 Hierarchy of Evidence (strongest to weakest)
- Systematic review / Meta-analysis
- Randomised Controlled Trial (RCT)
- Cohort study
- Case-control study
- Cross-sectional study
- Ecological study
- Case series / Case report
- Expert opinion
3.2 Study Design Comparison Table
| Feature | Cross-sectional | Case-control | Cohort | RCT |
|---|---|---|---|---|
| Direction | Snapshot | Backward (retrospective) | Forward (prospective or retrospective) | Forward (prospective) |
| Measure | Prevalence | Odds ratio | Relative risk, incidence | Relative risk, NNT |
| Exposure & outcome | Measured simultaneously | Outcome known, exposure sought | Exposure known, outcome awaited | Exposure assigned, outcome awaited |
| Causation | Cannot establish | Suggests association | Stronger association | Strongest causation |
| Time | One point | Past exposure recalled | Long follow-up | Fixed duration |
| Cost | Low | Moderate | High | Very high |
| Best for | Prevalence estimation, generating hypotheses | Rare diseases | Rare exposures, incidence estimation | Testing interventions |
3.3 Cross-sectional Study
- Design: Measures exposure and outcome at the same point in time
- Measure: Prevalence (not incidence)
- Strengths: Quick, cheap, good for generating hypotheses, estimates disease burden
- Limitations: Cannot establish temporality (chicken-or-egg problem), prone to prevalence bias (captures long-duration cases)
- Psychiatry example: NMHS 2016, national prevalence of mental disorders in India
3.4 Case-Control Study
- Design: Start with cases (disease +) and controls (disease −), look backward for exposure
- Measure: Odds Ratio (OR) = (a×d) / (b×c) from a 2×2 table
- Strengths: Efficient for rare diseases, relatively quick, inexpensive
- Limitations: Recall bias (cases remember exposures differently), selection bias in choosing controls, cannot calculate incidence or prevalence directly
- Psychiatry example: Childhood trauma (exposure) in patients with BPD (cases) vs. healthy controls
Interpreting Odds Ratio:
- OR = 1 → no association
- OR > 1 → exposure increases odds of disease
- OR < 1 → exposure is protective
- OR = 3.5 → the odds of disease are 3.5 times higher in exposed vs. unexposed
3.5 Cohort Study
- Design: Start with exposed and unexposed groups, follow forward to see who develops outcome
- Types: Prospective (follow from now) or retrospective (use existing records)
- Measure: Relative Risk (RR) = incidence in exposed / incidence in unexposed
- Strengths: Can establish temporality, calculate incidence, study multiple outcomes of one exposure
- Limitations: Expensive, time-consuming, loss to follow-up, not efficient for rare diseases
- Psychiatry example: Framingham Heart Study, depressive symptoms and cardiovascular events
Interpreting Relative Risk:
- RR = 1 → no association
- RR > 1 → exposure increases risk
- RR < 1 → exposure is protective
- RR = 2.0 → exposed group has twice the risk
RR vs OR: In rare diseases (prevalence < 10%), OR approximates RR. In common diseases, OR overestimates RR.
3.6 Randomised Controlled Trial (RCT)
- Design: Participants randomly assigned to intervention or control
- Gold standard for establishing causation
Key Components:
- Strengths: Highest internal validity, can establish causation, minimises bias
- Limitations: Expensive, ethical constraints (can't randomise harmful exposures), Hawthorne effect, low external validity if inclusion criteria too strict
- Psychiatry example: STAR*D trial, sequential treatment strategies for depression
3.7 Ecological Study
- Design: Uses aggregate data (populations, not individuals)
- Measure: Correlation at population level
- Strengths: Uses existing data, hypothesis-generating, cheap
- Limitations: Ecological fallacy, cannot infer individual-level associations from group-level data
- Psychiatry example: Countries with higher lithium in drinking water have lower suicide rates
3.8 Case Series and Case Report
| Feature | Case Report | Case Series |
|---|---|---|
| N | 1 patient | Multiple patients (no control group) |
| Use | Novel presentations, rare side effects, new treatment ideas | Describe patterns, generate hypotheses |
| Strength | First signal of new phenomenon | Identifies emerging patterns |
| Limitation | No comparison, no generalisability | No controls, no causation |
| Example | First report of NMS with a new antipsychotic | Series of 20 patients with treatment-resistant depression who responded to ketamine |
4. STATISTICAL TESTS
4.1 Key Concepts First
4.2 The Decision Tree: Which Test When?
Step 1: What type of data?
| Data Type | Examples | Tests |
|---|---|---|
| Continuous (interval/ratio) | HAM-D scores, age, weight | t-test, ANOVA, Pearson r |
| Categorical (nominal/ordinal) | Diagnosis (yes/no), severity (mild/moderate/severe) | Chi-square, Fisher's exact |
Step 2: How many groups?
| Groups | Continuous (Parametric) | Continuous (Non-parametric) | Categorical |
|---|---|---|---|
| 2 groups, unpaired | Independent t-test | Mann-Whitney U | Chi-square |
| 2 groups, paired | Paired t-test | Wilcoxon signed-rank | McNemar's test |
| 3+ groups, unpaired | One-way ANOVA | Kruskal-Wallis | Chi-square |
| 3+ groups, paired | Repeated measures ANOVA | Friedman test | Cochran's Q |
Step 3: Parametric or non-parametric?
Use parametric tests when:
- Data is normally distributed
- Continuous data (interval/ratio)
- Adequate sample size (n > 30 per group as rough guide)
- Homogeneity of variance (Levene's test)
Use non-parametric tests when:
- Data is ordinal or skewed
- Small sample size
- Normality assumption violated
- Outliers present
4.3 Individual Tests: Details
t-test (Student's t-test)
Independent (unpaired) t-test:
- Use: Compare means of two independent groups
- Assumptions: Normal distribution, continuous data, homogeneity of variance, independent observations
- Example: Mean HAM-D score in drug group vs. placebo group
Paired t-test:
- Use: Compare means of same group at two time points or matched pairs
- Assumptions: Normal distribution of differences, continuous data
- Example: HAM-D score before and after 6 weeks of treatment in the same patients
- Key: Calculates difference for each pair, then tests if mean difference = 0
ANOVA (Analysis of Variance)
One-way ANOVA:
- Use: Compare means of 3+ independent groups on one factor
- Example: Mean depression scores across three groups: SSRI, SNRI, placebo
- Post-hoc: If significant, use Tukey's HSD or Bonferroni to identify which pairs differ
Two-way ANOVA:
- Use: Examine effects of two independent variables and their interaction
- Example: Effect of drug (SSRI vs placebo) AND gender (male vs female) on depression score
Repeated measures ANOVA:
- Use: Same subjects measured at 3+ time points
- Example: Depression scores at baseline, 4 weeks, 8 weeks, 12 weeks
- Assumption: Sphericity (Mauchly's test), if violated, use Greenhouse-Geisser correction
Chi-Square Test
Test of Independence:
- Use: Test association between two categorical variables
- Example: Is there an association between gender (male/female) and diagnosis (depression/anxiety)?
- Requirement: Expected cell count ≥ 5 in at least 80% of cells
Goodness of Fit:
- Use: Test whether observed frequencies match expected frequencies
- Example: Are psychiatric diagnoses equally distributed across blood groups?
If expected cell count < 5 → use Fisher's exact test
Non-Parametric Tests
Mann-Whitney U test:
- Non-parametric alternative to independent t-test
- Compares rank distributions of two independent groups
- Example: Comparing satisfaction scores (ordinal) between two hospitals
Wilcoxon signed-rank test:
- Non-parametric alternative to paired t-test
- Compares two related samples
- Example: Patient-reported severity (ordinal) before and after intervention
Kruskal-Wallis test:
- Non-parametric alternative to one-way ANOVA
- Compares three or more independent groups
- Example: Comparing therapy satisfaction across 4 treatment modalities
- If significant, use Dunn's test for post-hoc pairwise comparisons
Correlation
| Feature | Pearson r | Spearman ρ |
|---|---|---|
| Data type | Continuous, normally distributed | Ordinal or non-normal continuous |
| Measures | Linear association | Monotonic association |
| Range | −1 to +1 | −1 to +1 |
| Example | Correlation between HAM-D and BDI scores | Correlation between severity ranking and treatment response ranking |
Interpreting r:
- 0–0.3: weak
- 0.3–0.7: moderate
- 0.7–1.0: strong
Correlation = Causation. Two variables may correlate because of a shared confounder.
Fisher's Exact Test
- Used when Chi-square assumptions are violated (expected count < 5)
- Small sample sizes
- Calculates exact probability rather than approximation
- Example: Association between a rare genotype and treatment response in a sample of 15 patients
Regression
| Type | Use | Example |
|---|---|---|
| Linear regression | Predict continuous outcome from one or more predictors | Predicting HAM-D score from age, duration of illness, baseline score |
| Logistic regression | Predict binary outcome from one or more predictors | Predicting treatment response (yes/no) from baseline variables |
| Multiple regression | Multiple predictors for one outcome | Effect of CBT hours, medication adherence, and social support on depression score |
5. SENSITIVITY AND SPECIFICITY
5.1 The 2×2 Table
| Disease + | Disease − | ||
|---|---|---|---|
| Test + | a (True Positive) | b (False Positive) | a + b |
| Test − | c (False Negative) | d (True Negative) | c + d |
| a + c | b + d | N |
5.2 Formulae
| Measure | Formula | Meaning |
|---|---|---|
| Sensitivity | a / (a + c) | Proportion of actual positives correctly identified (true positive rate) |
| Specificity | d / (b + d) | Proportion of actual negatives correctly identified (true negative rate) |
| PPV | a / (a + b) | Probability that a positive test = true disease |
| NPV | d / (c + d) | Probability that a negative test = true absence |
| LR+ | Sensitivity / (1 − Specificity) | How much a positive test increases the odds of disease |
| LR− | (1 − Sensitivity) / Specificity | How much a negative test decreases the odds of disease |
5.3 Clinical Rules
- SnNOut = Sensitive test, Negative result rules OUT disease
- SpPIn = Specific test, Positive result rules IN disease
Why? A highly sensitive test catches almost all true cases → very few false negatives → if it's negative, you can be confident the disease is absent.
5.4 Effect of Prevalence on PPV and NPV
| Prevalence | PPV | NPV |
|---|---|---|
| High (common disease) | ↑ Higher | ↓ Lower |
| Low (rare disease) | ↓ Lower | ↑ Higher |
Clinical implication: Even a highly sensitive and specific test will have poor PPV in a low-prevalence population because the absolute number of false positives exceeds true positives.
Example: If PHQ-9 sensitivity = 90%, specificity = 85%, and depression prevalence = 2%:
- For every 1000 people: 20 have depression, 980 do not
- True positives = 18 (90% of 20)
- False positives = 147 (15% of 980)
- PPV = 18 / (18 + 147) = 10.9%, most positive screens are false positives
5.5 ROC Curve
- ROC = Receiver Operating Characteristic
- Plots sensitivity (y-axis) vs 1 − specificity (x-axis) at different cut-off points
- AUC (Area Under the Curve): 0.5 = no better than chance; 0.7–0.8 = acceptable; 0.8–0.9 = excellent; > 0.9 = outstanding
- Use: Compare diagnostic performance of different scales or select optimal cut-off
- Psychiatry example: Comparing PHQ-9 vs BDI-II as screening tools for depression, the one with higher AUC performs better across all thresholds
6. EVIDENCE-BASED MEDICINE (EBM)
6.1 Definition
EBM = the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients. Integrates:
- Best available evidence
- Clinical expertise
- Patient values and preferences
(David Sackett, 1996)
6.2 Five Steps of EBM (Mnemonic: 5 A's)
| Step | Action | Example |
|---|---|---|
| 1. Ask | Formulate a PICO question | In adults with MDD (P), does CBT + SSRI (I) compared to SSRI alone (C) improve remission rates (O)? |
| 2. Acquire | Search for best evidence | PubMed, Cochrane Library, NICE guidelines |
| 3. Appraise | Critically evaluate the evidence | Check study design, bias, applicability |
| 4. Apply | Integrate evidence with clinical expertise and patient preferences | Discuss options with patient considering their values |
| 5. Assess | Evaluate your practice and outcomes | Audit remission rates in your patients |
6.3 Levels of Evidence Pyramid
| Level | Study Type | Strength |
|---|---|---|
| 1a | Systematic review of RCTs | Strongest |
| 1b | Individual RCT with narrow CI | ↓ |
| 2a | Systematic review of cohort studies | ↓ |
| 2b | Individual cohort study | ↓ |
| 3a | Systematic review of case-control studies | ↓ |
| 3b | Individual case-control study | ↓ |
| 4 | Case series | ↓ |
| 5 | Expert opinion | Weakest |
6.4 NNT and NNH
| Measure | Formula | Interpretation |
|---|---|---|
| NNT (Number Needed to Treat) | 1 / ARR | Number of patients needed to treat with intervention for one additional good outcome |
| NNH (Number Needed to Harm) | 1 / ARI | Number of patients treated before one additional patient is harmed |
| ARR (Absolute Risk Reduction) | CER − EER | Difference in event rates between control and experimental groups |
- CER = Control Event Rate
- EER = Experimental Event Rate
Example: If relapse rate is 40% with placebo and 25% with drug:
- ARR = 0.40 − 0.25 = 0.15
- NNT = 1 / 0.15 = 6.7 → treat 7 patients with the drug to prevent one relapse
Clinical rule: Lower NNT = more effective treatment. Higher NNH = safer treatment. Compare NNT to NNH for risk-benefit.
7. META-ANALYSIS AND SYSTEMATIC REVIEW
7.1 Definitions
| Feature | Systematic Review | Meta-analysis |
|---|---|---|
| Definition | Structured, reproducible synthesis of all available evidence on a question | Statistical pooling of results from multiple studies into a single summary estimate |
| Statistics | Qualitative synthesis (narrative) | Quantitative synthesis (statistical) |
| Relationship | Every meta-analysis is within a systematic review | Not every systematic review includes a meta-analysis |
| When meta-analysis is NOT done | Studies too heterogeneous, different outcomes, insufficient data |
7.2 Steps in a Systematic Review / Meta-analysis
- Define research question (PICO)
- Protocol registration (PROSPERO)
- Comprehensive literature search (multiple databases: PubMed, Cochrane, Embase, PsycINFO)
- Study selection (inclusion/exclusion criteria, PRISMA flowchart)
- Data extraction (standardised forms)
- Quality assessment (Cochrane Risk of Bias tool for RCTs, Newcastle-Ottawa Scale for observational studies)
- Data synthesis (narrative or meta-analytic)
- Assess heterogeneity (I2, Q test)
- Assess publication bias (funnel plot, Egger's test)
- Report (PRISMA guidelines)
7.3 Forest Plot Interpretation
A forest plot displays:
- Each study: horizontal line (CI) with a square (point estimate). Square size = study weight.
- Summary diamond: combined effect estimate; width = CI of the summary
- Line of no effect: vertical line at OR/RR = 1 (or mean difference = 0)
How to read:
- If the CI of a study crosses the line of no effect → that study is non-significant
- If the summary diamond does NOT cross the line → overall result is significant
- Wider CI = less precise study (smaller sample)
- Larger square = greater weight in the analysis
7.4 Heterogeneity
| Measure | Definition | Interpretation |
|---|---|---|
| Cochrane Q | Chi-square test for heterogeneity | p < 0.10 → significant heterogeneity (liberal threshold because of low power) |
| I2 | Percentage of variability due to heterogeneity rather than chance | 0–25% = low, 25–50% = moderate, 50–75% = substantial, > 75% = considerable |
If heterogeneity is high:
- Use random-effects model (assumes true effect varies across studies)
- Perform subgroup analysis or meta-regression to explore sources
- Consider NOT pooling data
If heterogeneity is low:
- Fixed-effect model is appropriate (assumes one true effect)
7.5 Publication Bias
- Problem: Studies with significant results are more likely to be published → pooled estimate is biased
- Funnel plot: Scatter plot of effect size (x-axis) vs study precision/sample size (y-axis). Symmetric = no bias. Asymmetric = possible publication bias.
- Egger's test: Statistical test for funnel plot asymmetry
- Trim and fill method: Statistical adjustment for missing studies
- Solutions: Pre-registration of trials, searching grey literature, contacting authors
7.6 Network Meta-analysis (NMA)
- Definition: Extends traditional meta-analysis to compare multiple treatments simultaneously, even if they haven't been compared head-to-head
- Advantage: Ranks treatments by probability of being best
- Key concept: Uses both direct evidence (A vs B trials) and indirect evidence (if A vs C and B vs C exist, can estimate A vs B)
- Assumption: Transitivity, similar study populations across comparisons
- Psychiatry landmark: Cipriani et al. (2018), network meta-analysis of 21 antidepressants. Found amitriptyline, mirtazapine, and venlafaxine most effective; fluoxetine, escitalopram most acceptable.
8. RESEARCH METHODOLOGY
8.1 Research Question: PICO Framework
| Component | Meaning | Example |
|---|---|---|
| P | Population | Adults with treatment-resistant depression |
| I | Intervention | Ketamine infusion |
| C | Comparison | Midazolam (active placebo) |
| O | Outcome | Reduction in MADRS score at 24 hours |
8.2 Hypothesis Types
| Type | Definition | Example |
|---|---|---|
| Null (H0) | No difference or no association | There is no difference in HAM-D scores between drug and placebo |
| Alternative (H1) | There IS a difference or association | Drug group has lower HAM-D scores than placebo |
| Directional (one-tailed) | Specifies the direction of difference | Drug group will have LOWER scores (not just different) |
| Non-directional (two-tailed) | Difference exists but direction unspecified | Scores will DIFFER between groups |
8.3 Sampling Methods
Probability Sampling (every member has a known chance of selection)
| Method | How It Works | Advantage | Limitation |
|---|---|---|---|
| Simple random | Every individual has equal chance (lottery, random number table) | Unbiased, representative | Needs complete sampling frame |
| Stratified random | Divide population into strata (e.g., age groups), random sample from each | Ensures representation of subgroups | More complex, needs knowledge of strata |
| Cluster | Divide population into clusters (e.g., villages), randomly select entire clusters | Practical for large, dispersed populations | Higher sampling error |
| Systematic | Every kth individual from a list (e.g., every 5th patient) | Simple, quick | Bias if list has a periodic pattern |
Non-Probability Sampling (not everyone has a chance of selection)
| Method | How It Works | Use Case |
|---|---|---|
| Convenience | Whoever is available | Quick pilot studies |
| Purposive | Researcher selects based on characteristics | Qualitative research, expert panels |
| Snowball | Participants recruit other participants | Hard-to-reach populations (IV drug users, sex workers) |
| Quota | Researcher ensures certain proportions (like stratified but non-random) | Market research |
8.4 Sample Size Calculation
Why it matters: Too small → Type II error (miss real effects). Too large → wasteful, ethical concerns (unnecessary exposure).
Factors affecting sample size:
General formula concept: n = f(Zα, Zβ, σ, δ) where:
- Zα = Z-value for desired significance level (1.96 for α = 0.05)
- Zβ = Z-value for desired power (0.84 for power = 0.80)
- σ = standard deviation
- δ = minimum clinically meaningful difference
8.5 Bias Types
| Bias | Definition | Example | How to Minimise |
|---|---|---|---|
| Selection bias | Systematic differences in who enters the study | Only including patients from tertiary hospital | Random sampling, clear inclusion criteria |
| Information bias | Systematic errors in measuring exposure or outcome | Recall bias in case-control studies | Standardised instruments, blinding |
| Recall bias | Cases remember exposures differently than controls | Mothers of children with birth defects recall medication use more accurately | Prospective design, objective records |
| Observer bias | Researcher's expectations influence data collection | Unblinded rater scores drug group as improved | Blinding of assessors |
| Performance bias | Unequal care between groups | Intervention group gets more attention | Standardised protocols, blinding |
| Attrition bias | Differential loss to follow-up | Sicker patients drop out of drug arm → drug looks better | ITT analysis, minimise dropout |
| Publication bias | Positive results more likely published | Only significant antidepressant trials published | Trial registries, grey literature search |
| Confounding | Third variable associated with both exposure and outcome | Coffee drinkers have more lung cancer (confounder: smoking) | Randomisation, restriction, matching, stratification, multivariate analysis |
| Lead-time bias | Earlier detection inflates survival time without true benefit | Screening detected cancer 2 years earlier but patient dies at same time | Measure mortality, not survival from diagnosis |
| Berkson bias | Hospital-based studies create spurious associations | Studying depression and diabetes in hospitalised patients, both conditions increase hospitalisation probability | Community-based studies |
8.6 Confounding, Effect Modification, and Mediation
| Concept | Definition | How to Identify | How to Handle |
|---|---|---|---|
| Confounding | Third variable distorts the true association | Crude and adjusted estimates differ | Randomisation, stratification, multivariate analysis |
| Effect modification | Effect of exposure on outcome DIFFERS across levels of a third variable | Significant interaction term | Report stratum-specific estimates (do NOT adjust away) |
| Mediation | Third variable lies on the causal pathway between exposure and outcome | Exposure → Mediator → Outcome | Mediation analysis (Baron & Kenny steps, Sobel test) |
8.7 Odds Ratio: Detailed
From a 2×2 table:
| Disease + | Disease − | |
|---|---|---|
| Exposed | a | b |
| Unexposed | c | d |
- OR = (a × d) / (b × c)
- Used in case-control studies (cannot calculate RR because you start with known outcome numbers)
- In cohort studies, RR is preferred; OR overestimates RR when outcome is common
8.8 Qualitative Research Methods
| Method | Focus | Approach | Example |
|---|---|---|---|
| Phenomenology | Lived experience of a phenomenon | In-depth interviews, meaning analysis | Experience of living with schizophrenia |
| Grounded theory | Develop theory from data | Iterative coding, constant comparison, theoretical sampling | How families cope with a member's addiction |
| Ethnography | Culture and social behaviour | Immersive observation in a community | Mental health beliefs in a tribal community |
| Content analysis | Systematic categorisation of text | Coding themes from transcripts | Themes in suicide notes |
| Thematic analysis | Identifying patterns across data | 6-step Braun & Clarke method | Barriers to mental health treatment in rural India |
Qualitative rigour criteria (equivalent to quantitative validity/reliability):
- Credibility (≈ internal validity), prolonged engagement, triangulation, member checking
- Transferability (≈ external validity), thick description
- Dependability (≈ reliability), audit trail
- Confirmability (≈ objectivity), reflexivity
8.9 Factor Analysis
- Definition: Statistical technique that reduces a large number of variables to fewer underlying dimensions (factors)
- Types:
- Exploratory Factor Analysis (EFA): No prior assumption about factor structure, "let the data speak"
- Confirmatory Factor Analysis (CFA): Tests whether a hypothesised factor structure fits the data
- Key terms:
- Factor loading: Correlation between a variable and a factor (> 0.3 or > 0.4 is meaningful)
- Eigenvalue: Amount of variance explained by a factor (Kaiser criterion: retain factors with eigenvalue > 1)
- Scree plot: Graph of eigenvalues, retain factors before the "elbow"
- Rotation: Varimax (orthogonal) or oblimin (oblique), simplifies factor structure
- Psychiatry applications:
- Development of PANSS factor structure (positive, negative, general psychopathology)
- Validation of personality inventories (e.g., five-factor model)
- Identifying symptom dimensions in depression (cognitive, somatic, affective)
9. ADDITIONAL HIGH-YIELD CONCEPTS
9.1 Normal Distribution
- Bell-shaped, symmetric curve
- Mean = median = mode
- 68% within ±1 SD, 95% within ±2 SD, 99.7% within ±3 SD
- Many statistical tests assume normality
- Test with: Shapiro-Wilk test, Kolmogorov-Smirnov test, Q-Q plot
9.2 Bonferroni Correction
- When doing multiple comparisons, risk of Type I error increases
- Bonferroni: divide α by number of comparisons (e.g., 3 comparisons → α = 0.05/3 = 0.017)
- Conservative, reduces power
- Alternative: Holm-Bonferroni (step-down, less conservative)
9.3 Intention-to-Treat vs Per-Protocol
| Feature | ITT | Per-Protocol |
|---|---|---|
| Includes | ALL randomised participants | Only those who completed protocol |
| Preserves | Randomisation | |
| Bias | Conservative (underestimates effect) | May overestimate effect |
| Preferred for | Superiority trials (primary analysis) | Complements ITT; preferred for non-inferiority trials |
9.4 CONSORT, STROBE, PRISMA
| Guideline | For Which Study Type | Key Feature |
|---|---|---|
| CONSORT | RCTs | Flow diagram of participant progression |
| STROBE | Observational studies (cohort, case-control, cross-sectional) | 22-item checklist |
| PRISMA | Systematic reviews and meta-analyses | Flow diagram of study selection |
KEY TABLES SUMMARY
Master Test Selection Table
| Scenario | Parametric Test | Non-Parametric Alternative |
|---|---|---|
| 2 independent groups, continuous | Independent t-test | Mann-Whitney U |
| 2 paired groups, continuous | Paired t-test | Wilcoxon signed-rank |
| 3+ independent groups, continuous | One-way ANOVA | Kruskal-Wallis |
| 3+ paired groups, continuous | Repeated measures ANOVA | Friedman test |
| 2 categorical variables | Chi-square test of independence | Fisher's exact test |
| Correlation (continuous) | Pearson r | Spearman ρ |
| Predict continuous outcome | Linear regression | |
| Predict binary outcome | Logistic regression |
Cross-reference: D2 (Model Answers), D3 (Mnemonics), D4 (Comparisons), D6 (Quick Review)
Model Answers
Q1. Define validity. Types. Importance in psychiatry. [10 marks: 2+6+2]
Exam Strategy
Split strictly: 1 para definition (2 marks), 4 types with examples (6 marks, 1.5 each), brief importance section (2 marks). Use a table for types.
Model Answer
Definition (2 marks)
Validity refers to the degree to which a measurement instrument accurately measures the construct it is intended to measure. A valid test measures what it claims to measure. For example, a valid depression scale should measure depression, not general distress or anxiety.
Types of Validity (6 marks)
| Type | Definition | Example |
|---|---|---|
| Face validity | The instrument appears to measure the intended construct on superficial examination. Assessed by non-experts. | PHQ-9 "looks like" it measures depression, questions about mood, sleep, appetite. Weakest form of validity. |
| Content validity | The instrument covers the entire domain of the construct comprehensively. Assessed by expert panel. | HAM-D includes items on depressed mood, guilt, insomnia, psychomotor changes, somatic symptoms, and suicidality, covering the full domain of depression. |
| Criterion validity | The instrument correlates with an established external standard (gold standard). Two subtypes: (a) Concurrent, new test and criterion measured simultaneously (e.g., comparing a new depression scale with HAM-D in same patients). (b) Predictive, test predicts a future outcome (e.g., MMSE score predicting progression to dementia at 2-year follow-up). | |
| Construct validity | The instrument measures the theoretical construct. Two subtypes: (a) Convergent, correlates with measures of the same construct (BDI correlates with PHQ-9). (b) Discriminant, does not correlate with measures of different constructs (BDI does not correlate with mania rating scale). |
Importance in Psychiatry (2 marks)
- Psychiatry lacks biomarkers, diagnosis and outcome measurement rely heavily on rating scales and clinical interviews. Validity ensures these instruments accurately capture subjective constructs like depression, psychosis, or personality pathology.
- Treatment trials depend on valid outcome measures; invalid instruments lead to erroneous conclusions about drug efficacy.
- Classification systems (ICD, DSM) require construct validity of diagnostic categories.
- Forensic psychiatry: medicolegal opinions must be based on valid assessment tools to be admissible.
Q2. What is validity? Discuss types with examples. [10 marks: 3+7]: LONG ESSAY CANDIDATE
Exam Strategy
This is a long essay candidate, write 3+ pages. Expand each type with psychiatric examples. Include internal vs external validity as bonus types to demonstrate depth.
Model Answer
Definition of Validity (3 marks)
Validity is defined as the degree to which a test or measurement instrument accurately measures the construct it purports to measure. It addresses the fundamental question: "Is this test measuring what we think it is measuring?"
Validity is not a binary property but exists on a continuum. A test may be valid for one purpose but not another. For example, the MMSE is a valid screening tool for cognitive impairment but not a valid diagnostic tool for specific dementia subtypes.
The concept is central to psychiatry because, unlike other medical specialties, psychiatric diagnoses rely primarily on phenomenological assessment rather than laboratory investigations. The validity of our tools directly determines the quality of our clinical decisions.
Types of Validity (7 marks)
1. Face Validity
Face validity refers to whether an instrument appears, on the surface, to measure what it claims to measure. It is the weakest and most subjective form of validity, typically assessed by non-experts or the target population.
Example: The PHQ-9 has high face validity, questions about feeling down, sleep difficulties, poor appetite, and concentration problems are recognisable as depression symptoms to most patients and clinicians.
Limitation: Face validity can be misleading. A test may "look right" but actually measure a different construct. Conversely, some valid tests may lack face validity (e.g., projective tests like the Rorschach).
Relevance: Important for patient compliance, patients are more likely to complete a questionnaire that appears relevant to their problems.
2. Content Validity
Content validity refers to the degree to which the instrument comprehensively covers all aspects of the construct being measured. It is assessed by expert panels who evaluate whether items represent the full domain.
Example: The HAM-D (Hamilton Depression Rating Scale) has strong content validity because it includes items assessing depressed mood, feelings of guilt, suicidal ideation, insomnia (early, middle, late), work and activities, psychomotor retardation, psychomotor agitation, anxiety, somatic symptoms, loss of insight, and diurnal variation, covering the full syndromal domain of depression.
A scale that only measured mood and sleep would lack content validity for depression because it omits cognitive, psychomotor, and somatic dimensions.
Assessment method: Content Validity Index (CVI), proportion of experts rating each item as relevant.
3. Criterion Validity
Criterion validity is established when the instrument's scores correlate with an external criterion or gold standard. It has two subtypes:
(a) Concurrent validity: The test and criterion are measured at the same time. The new instrument is validated against an established measure.
Example: A newly developed Hindi depression scale is administered alongside the validated HAM-D to the same patients on the same day. If the correlation (r) is > 0.70, concurrent validity is established.
(b) Predictive validity: The test score predicts a future outcome.
Example: The AUDIT (Alcohol Use Disorders Identification Test) score at admission predicting development of alcohol withdrawal seizures during hospitalisation. A high AUDIT score predicting more severe withdrawal supports predictive validity.
Another example: GCS score predicting functional outcome at 6 months in traumatic brain injury.
4. Construct Validity
Construct validity examines whether the instrument truly measures the theoretical construct it claims to measure. This is the most rigorous and important form of validity. Two subtypes:
(a) Convergent validity: The instrument correlates positively with other measures of the same construct.
Example: The BDI-II shows high correlation (r = 0.85) with the PHQ-9, both measuring depression. This supports convergent validity of both instruments.
(b) Discriminant (divergent) validity: The instrument does NOT correlate with measures of theoretically unrelated constructs.
Example: The BDI-II should show low correlation with the YMRS (Young Mania Rating Scale). If it showed high correlation with a mania scale, its discriminant validity would be questionable, it might be measuring general distress rather than depression specifically.
Other methods: Known-groups validity (instrument distinguishes between groups known to differ, e.g., depression scale scores higher in clinically depressed patients than healthy controls), factorial validity (factor analysis confirms the hypothesised factor structure).
5. Internal Validity (bonus)
Internal validity refers to the degree to which the results of a study can be attributed to the intervention rather than confounding variables. It applies to study designs rather than instruments.
Example: A well-conducted double-blind RCT of an antidepressant has high internal validity because randomisation controls for confounders, and blinding prevents bias.
Threats: Selection bias, maturation, history, testing effects, attrition, regression to the mean.
6. External Validity (bonus)
External validity refers to the generalisability of study results to other populations, settings, and times.
Example: An RCT conducted exclusively in urban tertiary centres with highly selected patients may lack external validity for rural primary care settings.
Relevance to Indian psychiatry: Many psychiatric treatment guidelines are based on Western studies, their external validity for the Indian population (different pharmacogenomics, cultural factors, disease presentation) requires careful consideration.
Q3. Define reliability and validity. Significance. Types. [10 marks: 3+2+5]
Exam Strategy
Balanced coverage needed. Give equal weight to both concepts. Use a comparison table to demonstrate understanding.
Model Answer
Definitions (3 marks)
Reliability is the degree to which a measurement instrument produces consistent and reproducible results when applied repeatedly under similar conditions. A reliable instrument yields the same result each time it measures the same underlying construct, assuming the construct has not changed.
Validity is the degree to which a measurement instrument accurately measures the construct it is intended to measure. A valid instrument measures what it claims to measure.
The relationship between reliability and validity is asymmetric: reliability is a necessary but not sufficient condition for validity. A test can be reliable but not valid (consistently measuring the wrong thing), but a test cannot be valid without being reliable.
Significance (2 marks)
- In psychiatry, where diagnosis depends on clinical assessment and rating scales rather than laboratory tests, both reliability and validity determine the quality of clinical practice, research, and medicolegal work.
- Inter-rater reliability is critical for multisite clinical trials, if raters at different centres score the HAM-D differently, the study results are uninterpretable.
- Validity of diagnostic categories (e.g., DSM-5) directly impacts treatment selection, prognosis estimation, and research classification.
- Reliable and valid instruments are prerequisites for evidence-based practice, guidelines based on studies using unreliable instruments are fundamentally flawed.
Types (5 marks)
Types of Reliability:
| Type | What It Measures | Statistic | Example |
|---|---|---|---|
| Test-retest | Stability over time | Pearson r, ICC | Administering BPRS to same patients 2 weeks apart |
| Inter-rater | Agreement between raters | Cohen's kappa | Two psychiatrists rating same patient on PANSS |
| Split-half | Internal consistency (two halves) | Spearman-Brown corrected r | Odd-numbered items vs even-numbered items of BDI |
| Internal consistency | How well items measure same construct | Cronbach's alpha | All 21 items of BDI measuring depression (α > 0.7) |
Types of Validity:
| Type | What It Measures | Example |
|---|---|---|
| Face | Surface appearance | PHQ-9 looks like a depression questionnaire |
| Content | Domain coverage | HAM-D covers all depression dimensions |
| Criterion (concurrent + predictive) | Correlation with gold standard | New scale correlates with HAM-D (concurrent); AUDIT predicts withdrawal severity (predictive) |
| Construct (convergent + discriminant) | Theoretical construct accuracy | BDI correlates with PHQ-9 (convergent); BDI does not correlate with YMRS (discriminant) |
Q4. What is reliability? Types with examples. Internal and external validity. [10 marks: 2+4+4]
Exam Strategy
Heavy on reliability (6 marks total with definition). Separate section for internal and external validity with threats.
Model Answer
Definition (2 marks)
Reliability is the extent to which a measurement instrument yields consistent, reproducible, and dependable results across different occasions, raters, or forms. If the underlying construct has not changed, a reliable instrument will produce the same score on repeated measurement.
Types of Reliability with Examples (4 marks)
1. Test-retest reliability: Measures the stability of scores over time. The same instrument is administered to the same individuals at two different time points, and the correlation between scores is calculated.
Example: The BPRS is administered to 50 patients with schizophrenia at baseline and again 2 weeks later (assuming clinical stability). A Pearson correlation coefficient of r = 0.85 indicates good test-retest reliability.
Consideration: The time interval matters, too short risks practice effects; too long risks genuine change in the construct.
2. Inter-rater reliability: Measures agreement between two or more raters assessing the same subject. Quantified using Cohen's kappa (κ) for categorical data or Intraclass Correlation Coefficient (ICC) for continuous data.
Example: Two psychiatrists independently rate the same videotaped clinical interview using the PANSS. Cohen's kappa of 0.75 indicates substantial inter-rater agreement. This is critical for multisite clinical trials where different raters must produce comparable scores.
3. Split-half reliability: The test is divided into two halves (usually odd vs even items), and the correlation between halves is calculated. Corrected using the Spearman-Brown prophecy formula.
Example: The 21-item BDI is split into odd-numbered items and even-numbered items. A corrected correlation of r = 0.88 indicates good split-half reliability.
4. Internal consistency (Cronbach's alpha): Measures how well all items in a test measure the same underlying construct. Essentially an average of all possible split-half combinations.
Example: Cronbach's alpha of 0.89 for the PHQ-9 means the 9 items consistently measure the same construct (depression). An alpha > 0.70 is acceptable; > 0.90 may indicate item redundancy.
Internal and External Validity (4 marks)
Internal Validity
Internal validity is the degree to which a study's results accurately reflect the true relationship between the independent and dependent variables, free from confounding influences. It answers: "Did the intervention truly cause the observed effect?"
Threats to internal validity:
- Selection bias: Groups differ at baseline (mitigated by randomisation)
- History: External events affect outcome during the study period
- Maturation: Natural developmental changes in participants
- Regression to the mean: Extreme values at baseline tend to move toward the mean on retest
- Attrition: Differential dropout between groups
- Instrumentation: Changes in measurement tools or observers over time
- Testing effect: Repeated measurement itself affects performance
Example: A double-blind, randomised, placebo-controlled trial of a new antipsychotic has high internal validity, randomisation balances confounders, blinding prevents observer and performance bias, and placebo controls for non-specific effects.
External Validity
External validity is the degree to which study findings can be generalised to other populations, settings, and times.
Threats to external validity:
- Selection bias in recruitment: Strict inclusion criteria limit generalisability
- Hawthorne effect: Participants behave differently because they know they are being observed
- Setting specificity: Results from academic centres may not apply to primary care
- Volunteer bias: Volunteers may differ systematically from the general population
- Cultural specificity: Western-derived treatments or scales may not apply across cultures
Example: The STAR*D trial enrolled patients from both primary care and psychiatric clinics across the US, enhancing external validity. However, it excluded patients with bipolar disorder and active substance dependence, limiting generalisability to those populations.
The tension: RCTs maximise internal validity (controlled conditions) but may sacrifice external validity (strict criteria create unrepresentative samples). Pragmatic trials attempt to balance both.
Q5. What is psychiatric epidemiology? Uses. Types of designs. [10 marks: 2+4+4]
Exam Strategy
Define psychiatric epidemiology, list uses with Indian examples, then cover 4 major study designs.
Model Answer
Definition (2 marks)
Psychiatric epidemiology is the study of the distribution, determinants, and outcomes of mental disorders in defined populations. It applies epidemiological methods to understand the frequency (incidence, prevalence), risk factors, natural history, and burden of psychiatric illnesses, and to evaluate the effectiveness of interventions.
It differs from clinical psychiatry in that its unit of analysis is the population rather than the individual patient.
Uses of Psychiatric Epidemiology (4 marks)
- Estimation of disease burden: Determines prevalence and incidence of mental disorders. The National Mental Health Survey (NMHS) 2016 estimated the lifetime prevalence of mental disorders in India at 13.7%, with a treatment gap of 70–92%.
- Identification of risk and protective factors: Identifying that childhood adversity increases risk of adult depression (ACE studies), or that social support is protective against PTSD after trauma.
- Planning mental health services: Prevalence data guides resource allocation. India's District Mental Health Programme (DMHP) expansion was informed by epidemiological data on service gaps.
- Evaluation of interventions: Community-based studies evaluating the effectiveness of DMHP, or the MANAS trial evaluating collaborative care for depression in primary care in Goa.
- Understanding natural history: Longitudinal studies tracking the course of first-episode psychosis inform treatment duration guidelines.
- Policy development: WHO's Mental Health Atlas, Global Burden of Disease data, and NMHS data inform the Mental Healthcare Act 2017 and National Mental Health Policy.
Types of Study Designs (4 marks)
1. Cross-sectional (Prevalence study)
- Measures exposure and outcome simultaneously
- Provides prevalence estimates
- Example: NMHS 2016, surveyed 34,802 individuals to estimate prevalence of mental disorders
- Limitation: Cannot establish temporality
2. Case-control study
- Compares individuals with disease (cases) to those without (controls), looks backward for exposure
- Measure: Odds ratio
- Example: Study comparing rates of childhood sexual abuse in women with BPD (cases) versus women with other personality disorders (controls)
- Limitation: Recall bias, selection of appropriate controls
3. Cohort study
- Follows exposed and unexposed groups forward (prospective) or backward using records (retrospective)
- Measure: Relative risk, incidence rate
- Example: Dunedin Multidisciplinary Health and Development Study, followed a birth cohort to study how childhood factors predict adult mental health
- Limitation: Expensive, time-consuming, loss to follow-up
4. Randomised Controlled Trial
- Random assignment to intervention or control
- Gold standard for causation
- Example: MANAS trial (Patel et al., 2010), collaborative stepped care for depression in primary care in Goa
- Limitation: Ethical constraints, cost, limited external validity
Q6. Describe different types of epidemiological studies in Psychiatry. [10 marks]
Exam Strategy
Cover all major types. Use a comparison table. Add psychiatric examples for each.
Model Answer
Introduction
Epidemiological studies in psychiatry can be classified as observational (researcher observes without intervening) or experimental (researcher assigns an intervention). Each design has specific strengths, limitations, and appropriate applications.
Classification of Epidemiological Studies
A. Observational Studies
1. Descriptive Studies
(a) Case report/case series:
- Description of one or more patients with unusual presentation or outcome
- Hypothesis-generating
- Example: First reports of neuroleptic malignant syndrome; first case series of catatonia responding to lorazepam
- Limitation: No control group, no generalisability
(b) Cross-sectional (Prevalence study):
- Measures exposure and disease status simultaneously in a defined population
- Provides prevalence, not incidence
- Measure: Prevalence ratio
- Example: WHO World Mental Health Survey (WMHS); National Mental Health Survey of India 2016
- Strengths: Quick, inexpensive, useful for health service planning
- Limitations: Cannot establish causation (temporal ambiguity), prevalence bias (over-represents chronic conditions)
(c) Ecological study:
- Uses aggregate data for populations, not individuals
- Example: Countries with higher lithium in drinking water showing lower suicide rates
- Limitation: Ecological fallacy, cannot infer individual-level associations from group-level data
2. Analytical Observational Studies
(a) Case-control study:
- Starts with cases (disease+) and controls (disease−), compares past exposure
- Direction: Retrospective (outcome → exposure)
- Measure: Odds Ratio (OR) = ad/bc
- Example: Childhood trauma in patients with dissociative identity disorder vs. controls; cannabis use in first-episode psychosis vs. controls
- Strengths: Efficient for rare diseases, quick, inexpensive
- Limitations: Recall bias, selection bias, cannot calculate incidence
(b) Cohort study:
- Starts with exposed and unexposed groups, follows forward
- Direction: Prospective or retrospective
- Measure: Relative Risk (RR), Incidence Rate Ratio
- Example: Edinburgh Postnatal Depression Scale at 6 weeks predicting attachment outcomes at 1 year (prospective cohort); Cannabis use in adolescence and risk of psychosis (Dunedin study)
- Strengths: Establishes temporality, calculates incidence, multiple outcomes
- Limitations: Expensive, time-consuming, loss to follow-up, not efficient for rare outcomes
Comparison of Analytical Observational Designs:
| Feature | Case-Control | Cohort |
|---|---|---|
| Direction | Backward | Forward |
| Starts with | Disease status | Exposure status |
| Measure | Odds Ratio | Relative Risk |
| Best for | Rare diseases | Rare exposures |
| Bias risk | Recall bias | Loss to follow-up |
| Cost | Lower | Higher |
B. Experimental Studies
(a) Randomised Controlled Trial (RCT):
- Random assignment to intervention or control group
- Gold standard for establishing causation
- Features: Randomisation, blinding (single/double/triple), allocation concealment, ITT analysis
- Example: CATIE trial (comparing atypical antipsychotics); STAR*D trial (sequential treatment for depression)
- Strengths: Highest internal validity, controls confounders
- Limitations: Expensive, ethical constraints, may lack external validity
(b) Quasi-experimental study:
- Intervention without randomisation
- Example: Before-after study of suicide rates following media guidelines implementation
- Limitation: Lacks randomisation, more vulnerable to confounders
(c) Community trial:
- Intervention applied at the community level
- Example: MANAS trial, cluster-randomised trial of collaborative care in Goa primary health centres
Q7. Discuss Non-Parametric Tests. [10 marks]
Exam Strategy
Define, explain when to use, then describe individual tests with examples. Include comparison table.
Model Answer
Definition and Rationale
Non-parametric tests (also called distribution-free tests) are statistical tests that do not assume the data follows a specific probability distribution (typically the normal distribution). They are used when:
- Data is ordinal or ranked
- Data is not normally distributed (skewed)
- Sample size is small
- Outliers are present that would distort parametric results
- Assumptions of parametric tests (normality, homogeneity of variance) are violated
Non-parametric tests work on ranks rather than raw values, making them more robust to outliers and distribution violations.
Individual Non-Parametric Tests
1. Mann-Whitney U Test
- Non-parametric alternative to the independent samples t-test
- Compares two independent groups by ranking all observations
- Null hypothesis: The distributions of the two groups are identical
- Example: Comparing patient satisfaction scores (ordinal: 1–5 Likert scale) between two hospitals
- Psychiatry example: Comparing CGI-Severity scores between patients on clozapine vs. olanzapine
2. Wilcoxon Signed-Rank Test
- Non-parametric alternative to the paired t-test
- Compares two related samples (before-after, matched pairs)
- Null hypothesis: The median difference between paired observations is zero
- Example: Comparing patient-reported anxiety severity (ordinal scale) before and after a mindfulness intervention in the same patients
3. Kruskal-Wallis Test
- Non-parametric alternative to one-way ANOVA
- Compares three or more independent groups
- Null hypothesis: All groups have the same distribution
- Post-hoc: If significant, use Dunn's test for pairwise comparisons
- Example: Comparing quality of life scores across four treatment groups (SSRI, SNRI, TCA, placebo) when scores are not normally distributed
4. Friedman Test
- Non-parametric alternative to repeated measures ANOVA
- Compares three or more related groups (same subjects at multiple time points)
- Example: Comparing symptom severity scores at baseline, 4 weeks, and 8 weeks in the same group of patients
5. Spearman's Rank Correlation (ρ)
- Non-parametric alternative to Pearson's r
- Measures the strength and direction of monotonic relationship between two ranked variables
- Example: Correlation between severity ranking on PANSS and level of social functioning (ordinal categories)
6. Fisher's Exact Test
- Used when Chi-square assumptions are violated (expected cell count < 5)
- Calculates exact probability for 2×2 tables with small samples
- Example: Association between a rare genetic polymorphism and treatment response in a sample of 12 patients
7. Sign Test
- Simplest non-parametric test for paired data
- Only considers direction of change (positive or negative), not magnitude
- Less powerful than Wilcoxon signed-rank but requires fewer assumptions
- Example: Whether more patients improved than worsened after an intervention
Comparison Table: Parametric vs Non-Parametric Equivalents
| Parametric | Non-Parametric | Purpose |
|---|---|---|
| Independent t-test | Mann-Whitney U | 2 independent groups |
| Paired t-test | Wilcoxon signed-rank | 2 related groups |
| One-way ANOVA | Kruskal-Wallis | 3+ independent groups |
| Repeated measures ANOVA | Friedman | 3+ related groups |
| Pearson r | Spearman ρ | Correlation |
| Chi-square | Fisher's exact | Categorical data (small n) |
Advantages of Non-Parametric Tests:
- No distribution assumptions required
- Can handle ordinal data and ranked data
- Robust to outliers
- Valid for small sample sizes
Limitations:
- Less statistical power than parametric equivalents (when parametric assumptions are met)
- Cannot easily handle complex designs (interactions, covariates)
- May lose information by converting continuous data to ranks
Q8. Chi Square test. [10 marks]
Exam Strategy
Define, explain types, give formula, worked example, assumptions, and psychiatric application.
Model Answer
Definition
The Chi-square (χ2) test is a non-parametric statistical test used to examine the association between two categorical (nominal or ordinal) variables. It compares observed frequencies with expected frequencies under the null hypothesis of no association.
Types of Chi-Square Test
1. Chi-Square Test of Independence
- Tests whether two categorical variables are independent of each other
- Used with contingency tables (2×2 or larger)
- Null hypothesis: The two variables are independent (no association)
- Example: Is there an association between gender (male/female) and diagnosis (depression/anxiety)?
2. Chi-Square Goodness of Fit
- Tests whether observed frequencies in a single categorical variable match expected frequencies
- Null hypothesis: Observed distribution matches expected distribution
- Example: Are psychiatric diagnoses equally distributed across four blood groups?
Formula
χ2 = Σ [(O − E)2 / E]
Where: O = observed frequency, E = expected frequency
For a 2×2 table, E for each cell = (row total × column total) / grand total
Degrees of Freedom: df = (rows − 1) × (columns − 1). For a 2×2 table, df = 1.
Worked Example
A study examines whether gender is associated with treatment response:
| Responded | Did not respond | Total | |
|---|---|---|---|
| Male | 30 | 20 | 50 |
| Female | 40 | 10 | 50 |
| Total | 70 | 30 | 100 |
Expected frequencies (under H0):
- Male, responded: (50 × 70)/100 = 35
- Male, not responded: (50 × 30)/100 = 15
- Female, responded: (50 × 70)/100 = 35
- Female, not responded: (50 × 30)/100 = 15
χ2 = (30-35)2/35 + (20-15)2/15 + (40-35)2/35 + (10-15)2/15
χ2 = 0.71 + 1.67 + 0.71 + 1.67 = 4.76
With df = 1, critical value at α = 0.05 is 3.84. Since 4.76 > 3.84, p < 0.05. We reject H0, there is a significant association between gender and treatment response.
Assumptions and Conditions
- Data must be categorical (nominal or ordinal)
- Observations must be independent (each subject contributes to only one cell)
- Expected frequency ≥ 5 in at least 80% of cells
- No cell should have an expected frequency < 1
- If assumptions are violated → use Fisher's exact test
Yates' Correction for Continuity
- Applied to 2×2 tables to correct for overestimation of χ2
| - Formula: χ2 = Σ [( | O − E | − 0.5)2 / E] |
|---|
- Makes the test more conservative (reduces Type I error)
Psychiatric Applications
- Association between adherence (yes/no) and relapse (yes/no)
- Comparing proportion of responders across treatment groups
- Association between family history of mental illness and diagnosis
- Comparing categorical outcomes (improved/not improved) across different psychotherapy modalities
Q9. Paired t Test. [10 marks]
Exam Strategy
Define, explain when to use, assumptions, formula, worked example with psychiatric data, interpretation.
Model Answer
Definition
The paired t-test (also called dependent samples t-test) is a parametric statistical test used to compare the means of two related measurements from the same subjects or matched pairs. It tests whether the mean difference between paired observations is significantly different from zero.
When to Use
- Same group of subjects measured at two time points (before-after design)
- Matched pairs of subjects (matched on key variables like age, gender)
- Each observation in one group has a natural pairing with an observation in the other group
Assumptions
- Continuous data (interval or ratio scale)
- Normal distribution of differences (not the raw scores, the differences between pairs must be approximately normal)
- Random sampling from the population
- Paired observations, each subject has two measurements
- No significant outliers in the difference scores
Formula
t = d / (SD_d / √n)
Where:
- d = mean of the differences (d = X1 − X2 for each pair)
- SD_d = standard deviation of the differences
- n = number of pairs
- df = n − 1
Worked Example
A psychiatrist administers the HAM-D to 8 patients before and after 6 weeks of SSRI treatment:
| Patient | Pre-treatment | Post-treatment | Difference (d) |
|---|---|---|---|
| 1 | 24 | 18 | 6 |
| 2 | 20 | 14 | 6 |
| 3 | 28 | 22 | 6 |
| 4 | 22 | 20 | 2 |
| 5 | 26 | 16 | 10 |
| 6 | 18 | 12 | 6 |
| 7 | 30 | 24 | 6 |
| 8 | 24 | 16 | 8 |
Calculations:
- d = (6+6+6+2+10+6+6+8)/8 = 50/8 = 6.25
- SD_d = 2.25 (calculated from deviations)
- t = 6.25 / (2.25/√8) = 6.25 / 0.796 = 7.85
- df = 8 − 1 = 7
- Critical t at α = 0.05, df = 7 (two-tailed) = 2.365
Since 7.85 > 2.365, p < 0.05. There is a statistically significant reduction in HAM-D scores after 6 weeks of SSRI treatment.
Interpretation
- The mean reduction in HAM-D score was 6.25 points (95% CI: 4.37–8.13)
- This is both statistically significant (p < 0.001) and clinically meaningful (> 3-point change on HAM-D is considered the MCID)
- Effect size: Cohen's d = d/SD_d = 6.25/2.25 = 2.78 (very large effect)
Advantages Over Independent t-test
- Controls for individual differences (each subject serves as their own control)
- More powerful because it reduces error variance
- Requires smaller sample sizes
When NOT to Use
- If the differences are not normally distributed → use Wilcoxon signed-rank test
- If data is ordinal → use Wilcoxon signed-rank test
- If more than two time points → use repeated measures ANOVA
Psychiatric Applications
- Pre-post treatment studies (HAM-D before and after ECT, medication, or psychotherapy)
- Comparing two scales administered to the same patient
- Matched case-control studies (matching on age, gender, duration of illness)
Q10. Enumerate statistical tests. Describe t-test with example. [10 marks: 4+6]
Exam Strategy
First part: classify all tests in a table (4 marks). Second part: detailed t-test with both types and an example (6 marks).
Model Answer
Enumeration of Statistical Tests (4 marks)
| Category | Parametric | Non-Parametric |
|---|---|---|
| 2 independent groups | Independent t-test | Mann-Whitney U |
| 2 paired groups | Paired t-test | Wilcoxon signed-rank |
| 3+ independent groups | One-way ANOVA | Kruskal-Wallis |
| 3+ related groups | Repeated measures ANOVA | Friedman test |
| Correlation | Pearson r | Spearman ρ |
| Categorical variables | Chi-square, Fisher's exact | |
| Prediction (continuous outcome) | Linear regression | |
| Prediction (binary outcome) | Logistic regression |
Additional tests: Two-way ANOVA (two factors), ANCOVA (adjusting for covariates), MANOVA (multiple dependent variables), McNemar's test (paired categorical), Cochran's Q (3+ paired categorical).
The t-test, Detailed Description (6 marks)
The t-test is a parametric statistical test that compares the means of one or two groups to determine if there is a statistically significant difference. Developed by William Sealy Gosset (published under the pseudonym "Student" in 1908).
General Assumptions:
- Continuous data (interval or ratio)
- Normal distribution of data (or of differences for paired t-test)
- Homogeneity of variance for independent t-test (assessed by Levene's test)
- Independent observations
Type 1: Independent (Unpaired) t-test
Purpose: Compares means of two independent groups.
Formula: t = (X1 − X2) / √(Sp2 × (1/n1 + 1/n2))
Where Sp2 is the pooled variance.
Example: Comparing mean HAM-D scores between 30 patients receiving fluoxetine and 30 patients receiving placebo after 8 weeks:
- Fluoxetine group: Mean = 10.5, SD = 4.2
- Placebo group: Mean = 16.8, SD = 5.1
- t = −5.35, df = 58, p < 0.001
- Interpretation: Fluoxetine significantly reduces HAM-D scores compared to placebo
When to use Welch's t-test: If Levene's test is significant (unequal variances), use Welch's t-test which does not assume equal variances.
Type 2: Paired (Dependent) t-test
Purpose: Compares means of two related measurements (before-after, matched pairs).
Formula: t = d / (SD_d / √n)
Example: Measuring PANSS scores in 25 patients before and 4 weeks after starting risperidone:
- Mean baseline PANSS: 85.6 (SD 12.3)
- Mean 4-week PANSS: 62.4 (SD 14.1)
- Mean difference: 23.2 (SD 8.7)
- t = 23.2 / (8.7/√25) = 23.2 / 1.74 = 13.33
- df = 24, p < 0.001
- Interpretation: There is a statistically significant reduction in PANSS scores after 4 weeks of risperidone
Type 3: One-sample t-test
Purpose: Compares a sample mean to a known or hypothesised population mean.
Example: Testing whether the mean IQ of patients with treatment-resistant schizophrenia differs from the population mean of 100.
Reporting Format
"An independent samples t-test revealed a statistically significant difference in HAM-D scores between the fluoxetine group (M = 10.5, SD = 4.2) and the placebo group (M = 16.8, SD = 5.1), t(58) = −5.35, p < 0.001, Cohen's d = 1.35."
Q11. Enumerate statistical tests. Describe paired t-test. Two non-parametric tests. [10 marks: 4+2+4]
Exam Strategy
Quick enumeration table (4 marks). Concise paired t-test (2 marks). Two non-parametric tests in detail (4 marks).
Model Answer
Enumeration of Statistical Tests (4 marks)
See Q10 for complete table. Briefly:
Parametric tests: Independent t-test, paired t-test, one-way ANOVA, two-way ANOVA, repeated measures ANOVA, Pearson correlation, linear regression, logistic regression, ANCOVA.
Non-parametric tests: Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis, Friedman, Spearman correlation, Chi-square, Fisher's exact test, McNemar's test, Sign test, Cochran's Q.
Paired t-test (2 marks)
The paired t-test compares the means of two related measurements from the same subjects. It calculates the difference (d) for each pair, then tests whether the mean difference (d) is significantly different from zero.
Formula: t = d / (SD_d / √n), df = n − 1
Assumptions: Continuous data, normal distribution of differences, paired observations.
Example: Comparing BDI scores in 20 patients before and after 12 sessions of CBT. If mean difference = 8.5 (SD = 3.2), t = 8.5/(3.2/√20) = 11.87, p < 0.001, significant improvement after CBT.
Two Non-Parametric Tests (4 marks)
1. Mann-Whitney U Test (2 marks)
The Mann-Whitney U test is the non-parametric alternative to the independent samples t-test. It compares the distributions of two independent groups by ranking all observations from both groups together.
Procedure:
- Combine all observations and rank them from lowest to highest
- Sum the ranks for each group
- Calculate U statistic for each group
- Compare U to the critical value or calculate z-score for large samples
Assumptions: Ordinal or continuous data, independent observations, similar distribution shapes (tests whether one distribution is shifted relative to the other).
Example: Comparing satisfaction with care (5-point Likert scale: very dissatisfied to very satisfied) between patients in a community mental health centre (n = 25) and a tertiary psychiatry department (n = 30). Since Likert data is ordinal, the Mann-Whitney U test is appropriate rather than the t-test.
2. Wilcoxon Signed-Rank Test (2 marks)
The Wilcoxon signed-rank test is the non-parametric alternative to the paired t-test. It compares two related measurements by analysing the magnitude and direction of differences.
Procedure:
- Calculate the difference for each pair
- Rank the absolute differences (ignoring zeros)
- Assign the sign (+ or −) of the original difference to each rank
- Sum the positive ranks (W+) and negative ranks (W−)
- The test statistic is the smaller of W+ and W−
Assumptions: Ordinal or continuous data, paired observations, symmetric distribution of differences.
Example: Comparing self-rated anxiety on a 10-point visual analogue scale before and after a 20-minute relaxation exercise in 15 patients. Since the VAS data may not be normally distributed with a small sample, the Wilcoxon signed-rank test is more appropriate than a paired t-test.
Q12. Describe Factor Analysis and applications in psychiatric research. [10 marks]
Exam Strategy
Define, explain EFA vs CFA, key terms, procedure, then psychiatric applications (the bulk of the marks).
Model Answer
Definition
Factor analysis is a multivariate statistical technique that reduces a large number of observed variables into a smaller number of underlying latent factors (dimensions) that explain the patterns of correlations among the variables. It identifies clusters of intercorrelated variables that are driven by the same underlying construct.
Types
1. Exploratory Factor Analysis (EFA)
- No prior hypothesis about factor structure
- The data determines how many factors exist and which variables load on which factor
- Used during scale development
- "Let the data speak"
2. Confirmatory Factor Analysis (CFA)
- Tests a pre-specified factor structure
- Uses structural equation modelling (SEM) to assess model fit
- Used to validate an existing scale or theoretical model
- Reports fit indices: CFI, RMSEA, SRMR, TLI
Key Concepts
Procedure of EFA
- Assess suitability: KMO > 0.6, Bartlett's test significant
- Extract factors: Principal Component Analysis (PCA) or Principal Axis Factoring
- Determine number of factors: Eigenvalue > 1 (Kaiser), scree plot, parallel analysis
- Rotate factors: Varimax (orthogonal) or oblimin (oblique)
- Interpret: Examine factor loadings, name factors based on loading pattern
- Report: Variance explained, factor loadings, communalities
Applications in Psychiatric Research
1. Scale Development and Validation
- The PANSS (Positive and Negative Syndrome Scale): Factor analysis identified the three-factor model (positive symptoms, negative symptoms, general psychopathology), later analyses have proposed 5-factor models
- The BDI: Factor analysis reveals cognitive-affective and somatic factors of depression
- The PCL-5 (PTSD Checklist): CFA confirms the four-factor DSM-5 model (intrusion, avoidance, negative cognitions/mood, hyperarousal)
2. Symptom Dimension Research
- Schizophrenia: Factor analysis of symptoms identified 5 dimensions, positive, negative, cognitive, affective, and excitement/hostility, influencing classification and treatment approaches
- Depression: Identifying cognitive, somatic, and affective dimensions helps in understanding why SSRIs and TCAs may work differently on different symptom clusters
- OCD: Factor analysis identified 4 symptom dimensions, contamination/cleaning, symmetry/ordering, forbidden thoughts, hoarding, influencing treatment selection
3. Personality Research
- The Five-Factor Model (Big Five: Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) was derived through factor analysis of personality descriptors
- Alternative Model of Personality Disorders (AMPD) in DSM-5 Section III uses factor-analytically derived personality trait domains
4. Diagnostic Classification
- Factor analysis of symptom data has informed revisions to ICD and DSM
- The p-factor (general psychopathology factor), analogous to g-factor in intelligence, was identified through bifactor analysis of psychiatric symptoms, suggesting a common liability to all mental disorders
5. Treatment Research
- Identifying which symptom dimensions respond to specific treatments
- The MATRICS consensus: factor analysis of cognitive tests identified 7 cognitive domains impaired in schizophrenia (speed of processing, attention, working memory, verbal learning, visual learning, reasoning, social cognition)
Q13. What is EBM? Steps. Practice with example. [10 marks: 2+2+6]: LONG ESSAY CANDIDATE
Exam Strategy
Long essay. After definition and steps, walk through a COMPLETE clinical example through all 5 steps. This is what earns the 6 marks.
Model Answer
Definition (2 marks)
Evidence-Based Medicine (EBM) is defined as "the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients" (Sackett et al., 1996). It integrates three pillars:
- Best available research evidence
- Clinical expertise of the practitioner
- Patient values and preferences
EBM does not mean blindly following evidence, it means integrating evidence with clinical judgment and the individual patient's context.
Five Steps (2 marks)
- ASK, Formulate a clear clinical question (PICO format)
- ACQUIRE, Search for the best available evidence
- APPRAISE, Critically evaluate the evidence for validity and applicability
- APPLY, Integrate evidence with clinical expertise and patient preferences
- ASSESS, Evaluate the outcome and the process
Practice with a Detailed Clinical Example (6 marks)
Clinical Scenario:
A 35-year-old man with moderate major depressive disorder (MDD) has had partial response to fluoxetine 20 mg for 8 weeks. You are considering augmentation strategies.
Step 1: ASK, Formulate the PICO question
PICO question: "In adults with MDD who have partially responded to an SSRI, does augmentation with aripiprazole compared to lithium lead to higher remission rates?"
Step 2: ACQUIRE, Search for evidence
Search strategy:
- PubMed: ("major depressive disorder" OR "depression") AND ("aripiprazole" OR "augmentation") AND ("lithium") AND ("randomised controlled trial")
- Cochrane Library: Search for systematic reviews on antidepressant augmentation
- NICE/APA guidelines on treatment-resistant depression
Key findings:
- Nelson and Papakostas (2009): Meta-analysis of aripiprazole augmentation, NNT = 9 for remission
- STAR*D Level 3: Lithium augmentation showed modest benefit
- Network meta-analysis (Zhou et al., 2015): Atypical antipsychotic augmentation superior to lithium for remission
Step 3: APPRAISE, Critical appraisal
For the network meta-analysis:
- Design: Systematic review with network meta-analysis (Level 1a evidence)
- Internal validity: Comprehensive search strategy, multiple databases, risk of bias assessment conducted
- Heterogeneity: I2 = 45% (moderate), some heterogeneity exists
- Publication bias: Funnel plot showed mild asymmetry
- Effect size: Aripiprazole augmentation: OR for remission = 2.1 (95% CI: 1.5–3.0) vs placebo
- NNT: ~9 for remission, NNH for weight gain: ~6
- Applicability: Most studies were in Western populations; limited Indian data
- Limitations: Industry sponsorship in many aripiprazole trials, short duration (6–8 weeks)
Step 4: APPLY, Integrate evidence with clinical context
- Evidence says: Aripiprazole augmentation has modest efficacy (NNT = 9) with metabolic side effects
- Clinical expertise: In my experience, lithium augmentation is well-tolerated in younger patients and has anti-suicidal properties
- Patient factors: This patient has a family history of bipolar disorder → lithium may serve dual purpose (augmentation + mood stabilisation). He is concerned about weight gain → aripiprazole's metabolic profile is relevant.
- Decision: After discussing options with the patient, including benefits (lithium's anti-suicidal properties, potential mood stabiliser benefit given family history) and risks (thyroid, renal monitoring needed), the patient opts for lithium augmentation.
This illustrates that EBM does not dictate a single correct answer, the "best evidence" is integrated with clinical expertise and patient preferences.
Step 5: ASSESS, Evaluate outcomes
- After 8 weeks of lithium augmentation (serum level 0.6–0.8 mEq/L), the patient achieves remission (HAM-D = 5)
- Thyroid and renal function remain normal
- The clinician reflects: Was this the right decision? Would auditing outcomes across all augmentation patients in the clinic provide useful data?
- This feeds back into quality improvement and future clinical questions
Q14. What is EBM? Practice in India? Limitations. [10 marks]
Exam Strategy
Quick definition, then focus on Indian context (unique section) and limitations.
Model Answer
Definition
Evidence-Based Medicine (EBM) is the integration of the best available research evidence with clinical expertise and patient values in making clinical decisions (Sackett, 1996). It follows 5 steps: Ask, Acquire, Appraise, Apply, Assess.
Practice of EBM in India
Current status:
- EBM is increasingly taught in medical curricula following CBME (Competency-Based Medical Education) guidelines
- Indian Psychiatric Society (IPS) publishes clinical practice guidelines adapted from international evidence
- Institutions like PG exams, and PG exams conduct systematic reviews and contribute to the Cochrane Collaboration
Challenges in Indian context:
- Limited Indian evidence base: Most psychiatric RCTs are conducted in Western, educated, industrialised, rich, democratic (WEIRD) populations. Pharmacogenomic differences (e.g., CYP2D6 polymorphisms in South Asians), cultural factors, and disease presentation may differ.
- Access to evidence: Despite Sci-Hub usage, many clinicians lack institutional access to journals. Open-access initiatives (PubMed Central, IJPM) help but coverage is incomplete.
- Treatment gap: With a 70–92% treatment gap (NMHS 2016) and 0.3 psychiatrists per 100,000 population, implementing evidence-based protocols is constrained by workforce shortages.
- Cultural applicability: Psychotherapy protocols (CBT, DBT) developed in Western settings require cultural adaptation for Indian patients (e.g., family involvement, spiritual frameworks, idioms of distress).
- Resource constraints: Evidence-based treatments like clozapine require blood monitoring, not feasible in rural areas. Cost of newer medications limits implementation.
- Research quality: Many Indian psychiatric studies have methodological limitations, small sample sizes, single-centre designs, lack of randomisation.
Positive developments:
- MANAS trial (Patel et al., 2010), a landmark Indian RCT demonstrating effectiveness of collaborative care for depression in Goa
- National Mental Health Survey 2016, first nationally representative psychiatric epidemiological data
- DBT India initiative, SCARF (Schizophrenia Research Foundation) in Chennai
- SPIRIT (South Asian Hub for Advocacy, Research, and Education in Mental Health)
Limitations of EBM
- Evidence hierarchy bias: Over-reliance on RCTs devalues clinical experience, qualitative research, and patient narratives
- Publication bias: Published evidence is skewed toward positive results
- Lag time: Evidence takes years to translate into practice (17-year bench-to-bedside gap)
- Not all questions are answerable by RCTs: Ethical/practical constraints (e.g., randomising to childhood adversity)
- Individual variation: Evidence provides population-level estimates, the individual patient may respond differently
- Industry influence: Pharmaceutical industry funding can bias trial design, outcome selection, and reporting
- Cookbook medicine risk: EBM misapplied becomes rigid algorithm-following, ignoring clinical nuance
- Equity issues: Evidence generated from privileged populations may not serve marginalised communities
Q15. Discuss levels of evidence and role in EBM. [10 marks]
Exam Strategy
Present the full hierarchy, explain each level, then discuss their role in clinical decision-making.
Model Answer
Introduction
Levels of evidence refer to a hierarchical ranking system that grades the quality and reliability of evidence from clinical research. This hierarchy helps clinicians quickly assess the strength of evidence supporting a clinical decision.
Levels of Evidence Hierarchy
| Level | Study Type | Description |
|---|---|---|
| 1a | Systematic review of RCTs | Pooled analysis of multiple RCTs with meta-analysis. Strongest evidence. |
| 1b | Individual RCT with narrow CI | Single well-designed RCT with adequate power and narrow confidence intervals. |
| 1c | All-or-none studies | Dramatic effects where all patients previously died but now survive (e.g., insulin for DKA). |
| 2a | Systematic review of cohort studies | Pooled analysis of observational cohort studies. |
| 2b | Individual cohort study or low-quality RCT | Single cohort study or an RCT with methodological limitations. |
| 2c | Outcomes research | Ecological studies using large databases. |
| 3a | Systematic review of case-control studies | Pooled analysis of case-control studies. |
| 3b | Individual case-control study | Single well-designed case-control study. |
| 4 | Case series | Report of a series of patients without a control group. |
| 5 | Expert opinion | Opinion of respected authorities, based on clinical experience, descriptive studies, or reports of expert committees. |
The Evidence Pyramid (Visual Description)
The evidence pyramid has its base as expert opinion (broadest, weakest) and its apex as systematic reviews/meta-analyses (narrowest, strongest). Moving up the pyramid:
- Base: Expert opinion, large volume, weakest methodology
- Case series/reports, descriptive, no controls
- Case-control studies, analytical, retrospective
- Cohort studies, analytical, prospective, can establish temporality
- RCTs, experimental, can establish causation
- Apex: Systematic reviews/meta-analyses, synthesise all available evidence
Role of Levels of Evidence in EBM
1. Guiding treatment decisions
- When Level 1 evidence exists, it should form the foundation of clinical practice
- Example: Cochrane review supporting SSRIs for moderate-severe depression (Level 1a), should be first-line treatment
2. Developing clinical guidelines
- Guidelines grade their recommendations based on evidence levels
- NICE, APA, WFSBP, and IPS guidelines use evidence levels to justify recommendations
- Grade A recommendation (based on Level 1 evidence) carries more weight than Grade C (based on Level 4-5)
3. Identifying evidence gaps
- When only Level 4-5 evidence exists for a clinical question, this highlights the need for RCTs
- Example: Ketamine for treatment-resistant depression initially had only case series (Level 4); this drove the design of RCTs
4. Clinical decision-making in practice
- Level 1 evidence may not always be available or applicable
- In such cases, clinicians use the best available evidence even if it is lower-level
- Example: For ultra-treatment-resistant depression, combination strategies often rely on Level 4-5 evidence because RCTs are lacking
5. Limitations of the hierarchy
- Higher level = always better for all questions
- Qualitative research (not on the hierarchy) is essential for understanding patient experience
- Well-designed observational studies can sometimes provide stronger evidence than flawed RCTs
- Network meta-analyses depend on the quality of included RCTs
- The hierarchy works best for treatment questions; questions about prognosis, diagnosis, and harm may need different frameworks
Q16. What is quantitative research? Features, strengths, limitations. [10 marks: 3+3+2+2]
Exam Strategy
Define with features (3+3 marks), then strengths and limitations (2+2 marks).
Model Answer
Definition (3 marks)
Quantitative research is a systematic empirical investigation that uses numerical data, statistical analysis, and mathematical models to test hypotheses, establish relationships between variables, and make generalisable inferences about populations. It follows the scientific method of hypothesis formulation, data collection, statistical testing, and interpretation.
The underlying philosophy is positivism, the belief that reality is objective, measurable, and can be understood through systematic observation and experimentation.
Features (3 marks)
- Numerical data: Variables are measured and expressed as numbers (e.g., HAM-D scores, age, number of hospitalisations)
- Hypothesis-driven: Begins with a clear, testable hypothesis (H0 and H1)
- Structured design: Follows predetermined protocols, study design, sampling, data collection, and analysis are planned before data collection begins
- Large sample sizes: Aims for adequate statistical power through sample size calculation
- Standardised instruments: Uses validated rating scales, structured interviews, laboratory measures
- Statistical analysis: Uses inferential statistics (t-test, ANOVA, regression) to test hypotheses
- Objectivity: Researcher maintains distance from participants; findings should be reproducible
- Generalisability: Aims to generalise findings from sample to population
- Control of variables: Attempts to control or adjust for confounders
- Replicability: Detailed methodology allows other researchers to replicate the study
Strengths (2 marks)
- Objectivity and precision: Numerical data minimises subjective interpretation
- Generalisability: Large, representative samples allow population-level inferences
- Causal inference: Experimental designs (RCTs) can establish cause-and-effect
- Statistical power: Can detect small but meaningful effects
- Reproducibility: Standardised methods allow replication
- Policy influence: Quantitative evidence (NNT, effect sizes) directly informs treatment guidelines
Limitations (2 marks)
- Reductionism: Complex human experiences (e.g., suffering, meaning, therapeutic relationship) are reduced to numbers
- Context stripping: Decontextualises findings, a mean HAM-D score doesn't capture the lived experience of depression
- Assumes measurability: Not all psychiatric constructs are easily quantifiable (e.g., therapeutic alliance, insight)
- Researcher bias in design: Choice of variables, outcome measures, and statistical tests involves subjective decisions
- Ecological validity: Controlled laboratory or clinical settings may not reflect real-world conditions
- Missing "why": Tells you THAT something works but not WHY or HOW patients experience it
Q17. Qualitative vs quantitative research. [10 marks: 2+5+3]
Exam Strategy
Define both (2 marks), detailed comparison table (5 marks), when to use each in psychiatry (3 marks).
Model Answer
Definitions (2 marks)
Quantitative research uses numerical data and statistical analysis to test hypotheses, measure variables, and generalise findings to populations. It follows a deductive approach (theory → hypothesis → observation).
Qualitative research uses non-numerical data (words, themes, narratives) to explore meaning, experience, and social processes. It follows an inductive approach (observation → patterns → theory).
Detailed Comparison (5 marks)
| Feature | Quantitative | Qualitative |
|---|---|---|
| Philosophy | Positivism (objective reality) | Constructivism/interpretivism (subjective reality) |
| Approach | Deductive (theory-driven) | Inductive (data-driven) |
| Data type | Numbers, measurements | Words, narratives, images |
| Sample size | Large (powered for statistics) | Small (purposive, data saturation) |
| Sampling | Random/probability | Purposive, snowball, theoretical |
| Analysis | Statistical (t-test, ANOVA, regression) | Thematic, content, grounded theory, IPA |
| Instruments | Standardised scales, questionnaires | Interviews, focus groups, observation |
| Researcher role | Detached, objective | Embedded, reflexive |
| Outcome | Generalisable findings, effect sizes | Rich descriptions, themes, theory |
| Question type | "How much? How many? Is there a difference?" | "What is the experience like? How do people make sense of it?" |
| Rigour criteria | Validity, reliability | Credibility, transferability, dependability, confirmability |
| Causation | Can establish (RCTs) | Cannot establish |
| Replication | High (standardised methods) | Low (context-dependent) |
When to Use Each in Psychiatry (3 marks)
Quantitative is preferred when:
- Testing drug efficacy (RCTs)
- Measuring prevalence and incidence (epidemiological surveys)
- Comparing treatment outcomes across groups
- Validating rating scales (psychometric studies)
- Example: "Does aripiprazole augmentation improve remission rates in treatment-resistant depression?", requires an RCT
Qualitative is preferred when:
- Exploring lived experience of illness (phenomenology)
- Understanding barriers to treatment adherence
- Developing culturally sensitive interventions
- Generating new hypotheses from patient narratives
- Example: "How do patients with schizophrenia experience auditory hallucinations?", requires phenomenological interviews
Mixed methods, integrating both:
- Increasingly used in psychiatric research
- Example: MANAS trial used quantitative outcomes (PHQ-9 remission) AND qualitative interviews with patients and health workers to understand implementation
- Sequential explanatory: Quantitative data first → qualitative to explain findings
- Sequential exploratory: Qualitative first → develop instrument → quantitative testing
Q18. What is meta-analysis vs systematic review? Key studies in psychiatry. [10 marks]
Exam Strategy
Define and compare (4 marks), process (3 marks), landmark psychiatric studies (3 marks).
Model Answer
Definitions and Comparison (4 marks)
| Feature | Systematic Review | Meta-analysis |
|---|---|---|
| Definition | A structured, reproducible method of identifying, evaluating, and synthesising all available evidence relevant to a specific research question | A statistical technique that combines quantitative results from multiple studies into a single pooled estimate |
| Nature | Qualitative synthesis | Quantitative synthesis |
| Relationship | May or may not include a meta-analysis | Always part of a systematic review |
| When meta-analysis is NOT done | When studies are too heterogeneous, use different outcome measures, or have insufficient data for pooling | |
| Output | Narrative synthesis with tables | Forest plot, pooled effect size, I2, funnel plot |
A meta-analysis is always nested within a systematic review, but not every systematic review performs a meta-analysis. When study heterogeneity is too high or studies are too clinically diverse, a narrative synthesis is more appropriate.
Process of Conducting a Systematic Review/Meta-analysis (3 marks)
- Formulate PICO question and register protocol (PROSPERO)
- Search strategy: Multiple databases (PubMed, Cochrane, Embase, PsycINFO), grey literature, hand-searching reference lists
- Study selection: Two independent reviewers screen titles/abstracts, then full texts; PRISMA flow diagram documents this
- Data extraction: Standardised forms extracting sample size, effect sizes, outcomes
- Quality assessment: Cochrane Risk of Bias tool (RCTs), Newcastle-Ottawa Scale (observational)
- Data synthesis:
- Fixed-effect model (assumes one true effect across studies)
- Random-effects model (assumes true effect varies, used when heterogeneity exists)
- Assess heterogeneity: I2 statistic, Cochrane Q test
- Assess publication bias: Funnel plot, Egger's test, trim-and-fill
- Sensitivity analysis: Removing one study at a time to test robustness
- Report: Following PRISMA guidelines
Forest plot interpretation: Each study represented by a square (point estimate) with horizontal line (CI). Diamond at bottom = pooled estimate. Vertical line of no effect (OR = 1 or MD = 0). If diamond does not cross this line → overall result is statistically significant.
Key Meta-analyses and Systematic Reviews in Psychiatry (3 marks)
- Cipriani et al., Lancet 2018: Network meta-analysis of 21 antidepressants (522 trials, 116,477 patients). Found all antidepressants more effective than placebo. Amitriptyline, mirtazapine, and venlafaxine most effective; fluoxetine and escitalopram best accepted (lowest dropout).
- Leucht et al., Lancet 2012: Meta-analysis of antipsychotics for schizophrenia (65 RCTs). Found all antipsychotics more effective than placebo; effect sizes were moderate. Clozapine, amisulpride, olanzapine, and risperidone showed largest effects.
- Cuijpers et al. (multiple reviews): Extensive meta-analyses of psychotherapy for depression. CBT, IPT, and behavioural activation all effective; no consistent superiority of one modality over others (Dodo bird verdict). Combined psychotherapy + medication superior to either alone.
- Furukawa et al., World Psychiatry 2019: Network meta-analysis of initial treatment strategies for depression (pharmacotherapy vs psychotherapy vs combined). Combined treatment most effective.
- NICE guidelines for schizophrenia, depression, bipolar disorder: Based on systematic reviews of the evidence.
- Cochrane reviews in psychiatry: Cover topics from ECT for depression to clozapine for treatment-resistant schizophrenia to psychoeducation for bipolar disorder.
Q19. Odds ratio. Confounding, effect modification, mediation. [10 marks: 4+2+2+2]
Exam Strategy
Detailed OR section with 2×2 table (4 marks). Then concise but clear explanations of each concept (2 marks each).
Model Answer
Odds Ratio (4 marks)
Definition: The odds ratio (OR) is a measure of association between an exposure and an outcome in case-control and cross-sectional studies. It compares the odds of exposure in cases to the odds of exposure in controls.
2×2 Table:
| Disease + (Cases) | Disease − (Controls) | |
|---|---|---|
| Exposed | a | b |
| Unexposed | c | d |
Formula: OR = (a × d) / (b × c)
Interpretation:
- OR = 1 → No association between exposure and outcome
- OR > 1 → Exposure is associated with increased odds of disease (risk factor)
- OR < 1 → Exposure is associated with decreased odds of disease (protective factor)
- The 95% CI should not include 1 for the OR to be statistically significant
Example: A case-control study of childhood trauma and adult depression:
| Depression (Cases) | No Depression (Controls) | |
|---|---|---|
| Trauma + | 80 | 40 |
| Trauma − | 20 | 60 |
OR = (80 × 60) / (40 × 20) = 4800/800 = 6.0
Interpretation: Individuals with childhood trauma have 6 times the odds of developing depression compared to those without trauma.
OR vs RR: In case-control studies, we cannot calculate relative risk (because we select by outcome, not by exposure). OR approximates RR when the outcome is rare (< 10%).
Confounding (2 marks)
Definition: Confounding occurs when a third variable (confounder) is associated with both the exposure and the outcome, creating a spurious or distorted association between them.
Criteria for a confounder:
- Associated with the exposure
- Independently associated with the outcome
- Not on the causal pathway between exposure and outcome
Example: A study finds that coffee consumption is associated with lung cancer. However, coffee drinkers are more likely to smoke. Smoking is the confounder, it is associated with both coffee consumption and lung cancer. Once you adjust for smoking, the coffee-lung cancer association disappears.
How to control confounding:
- At design stage: Randomisation (best), restriction, matching
- At analysis stage: Stratification, multivariate regression, propensity score matching
Detection: If the adjusted (stratified or regression-adjusted) estimate differs from the crude estimate by > 10%, confounding is present.
Effect Modification (Interaction) (2 marks)
Definition: Effect modification occurs when the magnitude of the association between an exposure and an outcome differs across levels of a third variable (the effect modifier). Unlike confounding, effect modification is a real biological or social phenomenon and should be reported, not adjusted away.
Example: A study finds that SSRI treatment reduces depression scores by 8 points in women but only 3 points in men. Gender is an effect modifier, the treatment effect differs by gender.
How to detect:
- Stratified analysis: different effect sizes in different strata
- Regression: significant interaction term (exposure × modifier)
How to handle: Report stratum-specific estimates. Do NOT adjust for effect modifiers (unlike confounders). If SSRIs work better in women, this is clinically important information.
Mediation (2 marks)
Definition: Mediation occurs when a third variable (mediator) lies on the causal pathway between the exposure and the outcome. The exposure causes the mediator, which in turn causes the outcome.
Causal pathway: Exposure → Mediator → Outcome
Example: Childhood trauma → Maladaptive schemas → Adult depression. Maladaptive schemas mediate the relationship between trauma and depression. Part of the effect of trauma on depression operates through the development of maladaptive schemas.
Testing mediation (Baron & Kenny, 1986):
- Exposure significantly predicts outcome (path c)
- Exposure significantly predicts mediator (path a)
- Mediator significantly predicts outcome when controlling for exposure (path b)
- The effect of exposure on outcome is reduced (partial mediation) or non-significant (full mediation) when mediator is in the model (path c')
Types:
- Full mediation: Exposure effect becomes non-significant after including mediator
- Partial mediation: Exposure effect is reduced but remains significant
Q20. Sampling techniques. Sample size. Principles of calculation. [10 marks: 2+2+6]
Exam Strategy
Quick sampling overview (2 marks), define sample size (2 marks), then detailed principles of calculation (6 marks).
Model Answer
Sampling Techniques (2 marks)
Probability sampling (every member has a known chance of selection):
- Simple random: Equal chance for all (lottery, random number table)
- Stratified: Divide into strata, sample from each (ensures subgroup representation)
- Cluster: Randomly select entire clusters (villages, schools)
- Systematic: Every kth individual from a list
Non-probability sampling (not everyone has equal chance):
- Convenience: Whoever is available
- Purposive: Selected by researcher for specific characteristics
- Snowball: Participants recruit others (hard-to-reach populations)
Sample Size (2 marks)
Sample size is the number of participants required in a study to detect a clinically meaningful effect with adequate statistical power while maintaining a specified level of significance. An adequate sample size is the ethical minimum, too small wastes resources and cannot answer the question, too large exposes unnecessary participants to experimental conditions.
Principles of Sample Size Calculation (6 marks)
1. Type I Error Rate (α)
- The probability of rejecting the null hypothesis when it is true (false positive)
- Conventionally set at 0.05 (5%)
- Corresponds to Zα = 1.96 for two-tailed tests
- Lowering α (e.g., to 0.01) increases the required sample size
- Zα for one-tailed test at 0.05 = 1.645
2. Statistical Power (1 − β)
- The probability of correctly rejecting the null hypothesis when it is false
- Conventionally set at 0.80 (80%), sometimes 0.90
- β = Type II error rate (false negative)
- Zβ = 0.84 for 80% power; 1.28 for 90% power
- Higher power → larger sample required
3. Effect Size (δ)
- The minimum clinically meaningful difference between groups
- Smaller expected effect sizes require larger samples
- Determined by clinical judgment and prior research
- Cohen's conventions: small = 0.2, medium = 0.5, large = 0.8
4. Variability (σ)
- Standard deviation of the outcome measure
- Greater variability → larger sample needed
- Estimated from pilot data or previous studies
5. Sample Size Formulas
For comparing two means (independent t-test):
n (per group) = 2 × [(Zα + Zβ)2 × σ2] / δ2
Where:
- Zα = 1.96 (for α = 0.05, two-tailed)
- Zβ = 0.84 (for power = 0.80)
- σ = standard deviation
- δ = minimum clinically meaningful difference
Worked example:
Comparing HAM-D scores between drug and placebo groups:
- Expected difference (δ) = 4 points
- SD (σ) = 6 points (from previous literature)
- α = 0.05, power = 0.80
n = 2 × [(1.96 + 0.84)2 × 36] / 16 = 2 × [7.84 × 36] / 16 = 2 × 17.64 = 35.3
Need approximately 36 per group, 72 total.
For comparing two proportions:
n = [(Zα√(2pq) + Zβ√(p1q1 + p2q2))2] / (p1 − p2)2
6. Adjustment for Dropout
Adjusted n = n / (1 − expected dropout rate)
If expecting 20% dropout: 72 / 0.80 = 90 total
7. Other Considerations
- Study design: Cluster randomised trials need inflation factor for intra-cluster correlation
- Multiple comparisons: Bonferroni or other corrections require larger samples
- Non-inferiority trials: Require larger samples than superiority trials (must rule out pre-specified non-inferiority margin)
- Crossover designs: Generally require smaller samples than parallel designs (within-subject comparison)
- Rare events: Very large samples needed (or alternative designs like case-control)
8. Software and Resources
- G*Power (free software for sample size calculation)
- nQuery, PASS (commercial)
- Online calculators (OpenEpi, Epitools)
- Consulting a biostatistician is standard practice for grant applications
Cross-reference: D1 (Study Notes), D3 (Mnemonics), D4 (Comparisons), D6 (Quick Review)
Mnemonics & Memory Tricks
Mnemonic 1: Types of Validity: "FaCe CoCo"
Fa = Face validity
Ce = Content validity (Experts judge domain coverage)
Co = Criterion validity (Concurrent + Predictive, against a gold standard)
Co = Construct validity (Convergent + Discriminant, theoretical construct)
Memory aid: "FaCe CoCo", like the face of a coconut shell. Face is surface level (weakest), and you crack through to get to the real content, criterion, and construct inside.
Expanded:
- Face → Looks right (subjective, weakest)
- Content → Covers everything (experts check the domain)
- Criterion → Correlates with gold standard (Concurrent = same time; Predictive = future)
- Construct → Measures the right concept (Convergent = correlates with same; Discriminant = doesn't correlate with different)
Mnemonic 2: Types of Reliability: "TISI"
T = Test-retest (Time stability)
I = Inter-rater (Individuals agree)
S = Split-half (Splitting the test)
I = Internal consistency (Items hang together, Cronbach's alpha)
Memory aid: "TISI" sounds like "Tissue", reliability is like a tissue, consistent and uniform throughout. If one part tears easily while the rest is strong, the tissue is not reliable.
Statistics to remember:
- Test-retest → Pearson r or ICC
- Inter-rater → Cohen's kappa
- Split-half → Spearman-Brown corrected r
- Internal consistency → Cronbach's alpha
Mnemonic 3: Study Design Hierarchy: "Smart Researchers Can Create Excellent Cases"
From strongest to weakest:
Systematic review / Meta-analysis
Randomised Controlled Trial
Cohort study
Case-control study
Epidemiological cross-sectional study
Case series / Case report
(Expert opinion at the bottom)
Memory aid: "Smart Researchers Can Create Excellent Cases", and the smartest ones (systematic reviews) sit at the top.
Mnemonic 4: Parametric ↔ Non-Parametric Test Pairs: "I Must Pay Willingly, One Keeps Repeating Freely, Plus Spares"
| Parametric | Non-Parametric | Mnemonic Word |
|---|---|---|
| Independent t-test | Mann-Whitney U | "I Must" |
| Paired t-test | Wilcoxon signed-rank | "Pay Willingly" |
| One-way ANOVA | Kruskal-Wallis | "One Keeps" |
| Repeated measures ANOVA | Friedman | "Repeating Freely" |
| Pearson r | Spearman ρ | "Plus Spares" |
Memory aid: "I Must Pay Willingly, One Keeps Repeating Freely, Plus Spares", imagine paying your statistician willingly because they keep repeating analyses freely and always have spare tests.
Mnemonic 5: Five Steps of EBM: "5 A's"
Ask → Formulate PICO question
Acquire → Search for evidence
Appraise → Critically evaluate evidence
Apply → Integrate with expertise and patient values
Assess → Evaluate outcomes
Memory aid: Already a built-in mnemonic. The 5 A's. Think "AAAAA", like the five-star rating you're giving to your evidence-based practice.
Mnemonic 6: Levels of Evidence: "Meta Rules Cohorts, Cases Come Last, Experts Guess"
Meta-analysis / Systematic review (Level 1a)
Randomised Controlled Trial (Level 1b)
Cohort study (Level 2)
Case-control study (Level 3)
Case series (Level 4)
Last: Expert opinion (Level 5)
Guess = expert opinion is basically an educated guess
Memory aid: "Meta Rules, Cohorts and Cases Come Last, Experts Guess", meta-analyses rule the hierarchy, and experts at the bottom are basically guessing (educated guessing, but still).
Mnemonic 7: Sensitivity/Specificity: "SnNOut / SpPIn"
SnNOut = Sensitive test, Negative result, rules Out disease
SpPIn = Specific test, Positive result, rules In disease
Memory aid: "SnNOut", if you're sick and the sensitive test says No, you're Out (no disease). "SpPIn", the Specific test says Positive, you're In (have the disease).
Why this works:
- High sensitivity → very few false negatives → negative result is trustworthy
- High specificity → very few false positives → positive result is trustworthy
Bonus, the 2×2 table positions:
- Sensitivity = a/(a+c), left column (disease + only)
- Specificity = d/(b+d), right column (disease − only)
- PPV = a/(a+b), top row (test + only)
- NPV = d/(c+d), bottom row (test − only)
Mnemonic 8: Bias Types: "SCORES + PB"
Selection bias, who enters the study
Confounding, third variable distorts
Observer bias, researcher sees what they expect
Recall bias, cases remember differently
Equipment/Instrumentation bias, tools change
Survivor/Attrition bias, differential dropout
+PB:
Performance bias, unequal care between groups
Berkson bias, hospital-based studies create false associations
Memory aid: "SCORES + PB", think of "SCORES" like test scores that can be biased in many ways, plus "PB" for publication bias (which you should always mention in any meta-analysis answer).
Mnemonic 9: PICO Framework: "Patient, Intervention, Comparison, Outcome"
P = Patient/Population/Problem
I = Intervention (or Exposure)
C = Comparison (or Control)
O = Outcome
Memory aid: "PICO" already sounds like "pick-o", you PICK your research question elements. For observational studies, some use PECO (E = Exposure).
Quick template: "In [P], does [I] compared to [C] improve [O]?"
Mnemonic 10: Type I vs Type II Errors: "The Boy Who Cried Wolf"
Type I error (α) = False alarm = "Crying wolf when there is no wolf"
- Rejecting H0 when it is TRUE
- Finding a significant result when there isn't one
- The p-value controls this (threshold: 0.05)
- Mnemonic: Type I = I made it up (false positive)
Type II error (β) = Missed wolf = "Not crying wolf when the wolf is real"
- Failing to reject H0 when it is FALSE
- Missing a real effect
- Power (1 − β) controls this
- Mnemonic: Type II = II stupid to miss it (false negative)
Memory aid: "Type I = I see it (but it's not there). Type II = Too blind to see it (but it IS there)."
Mnemonic 11: When to Use Which Test: "The 2-3 Rule"
2 groups?
- Independent → t-test (parametric) / Mann-Whitney (non-parametric)
- Paired → Paired t-test / Wilcoxon
3+ groups?
- Independent → ANOVA / Kruskal-Wallis
- Repeated → RM-ANOVA / Friedman
Categorical?
- 2×2 table → Chi-square (or Fisher if n < 5 expected)
- Paired categorical → McNemar's test
Correlation?
- Normal → Pearson
- Non-normal/ordinal → Spearman
Decision flowchart in words:
- What's the outcome? → Continuous or Categorical?
- If continuous: How many groups? → 2 or 3+?
- If 2 groups: Paired or Independent?
- Normally distributed? → Yes = parametric, No = non-parametric
Mnemonic 12: Threats to Internal Validity: "SHRIMP-T"
S = Selection bias
H = History (external events)
R = Regression to the mean
I = Instrumentation changes
M = Maturation (natural changes)
P = Practice/Testing effects
T = Testing attrition (dropout)
Memory aid: "A SHRIMP-T study has many internal validity threats", picture a tiny shrimp of a study, full of holes and threats. A well-designed RCT is the big fish that overcomes these.
Mnemonic 13: Cohen's Kappa Interpretation: "Slight Fair Moderate Substantial Perfect"
| Kappa | Agreement | Mnemonic |
|---|---|---|
| < 0.20 | Slight | Someone barely trying |
| 0.21–0.40 | Fair | Fair attempt |
| 0.41–0.60 | Moderate | Middle ground |
| 0.61–0.80 | Substantial | Solid work |
| 0.81–1.00 | Almost perfect | Amazing |
Memory aid: "Some Fair Midfielders Score Amazing goals", kappa goes from slight (barely trying) to almost perfect (amazing).
Mnemonic 14: Forest Plot Reading: "SLIDE"
S = Square = individual study point estimate (size = weight)
L = Line through square = confidence interval for that study
I = Invisible vertical line at null (OR=1 or MD=0) = line of no effect
D = Diamond at bottom = pooled/summary estimate (width = CI)
E = Effect is significant if diamond does NOT cross the line of no effect
Memory aid: "SLIDE", you slide your eyes down the forest plot from individual studies to the summary diamond.
Mnemonic 15: Qualitative Rigour: "CriTDeC"
Qualitative equivalents of quantitative concepts:
Cri = Credibility (≈ internal validity)
T = Transferability (≈ external validity)
De = Dependability (≈ reliability)
C = Confirmability (≈ objectivity)
Memory aid: "CriTDeC", say it fast, sounds like "critical deck", your critical deck of cards for evaluating qualitative research quality.
BONUS: Quick Number Anchors
Cross-reference: D1 (Study Notes), D2 (Model Answers), D4 (Comparisons), D6 (Quick Review)
High-Yield Comparisons
Table 1: Parametric vs Non-Parametric Tests (with Paired Equivalents)
| Purpose | Parametric Test | Non-Parametric Test | Data Type |
|---|---|---|---|
| 2 independent groups | Independent t-test | Mann-Whitney U | Continuous |
| 2 paired/related groups | Paired t-test | Wilcoxon signed-rank | Continuous |
| 3+ independent groups | One-way ANOVA | Kruskal-Wallis | Continuous |
| 3+ related groups | Repeated measures ANOVA | Friedman test | Continuous |
| Correlation | Pearson r | Spearman ρ | Continuous |
| 2×2 categorical | Chi-square / Fisher's exact | Categorical | |
| Paired categorical (2×2) | McNemar's test | Categorical | |
| 3+ paired categorical | Cochran's Q | Categorical |
When to use parametric: Normal distribution, continuous data, adequate sample size (n > 30 per group), homogeneity of variance.
When to use non-parametric: Ordinal data, skewed distribution, small sample size, outliers, violated assumptions.
Key distinction: Parametric tests use raw values and assume a specific distribution; non-parametric tests use ranks and are distribution-free.
Table 2: Type I vs Type II Errors
| Feature | Type I Error (α) | Type II Error (β) |
|---|---|---|
| Definition | Rejecting H0 when it is true | Failing to reject H0 when it is false |
| In plain language | Finding a difference when none exists | Missing a real difference |
| Analogy | False alarm (crying wolf) | Missed detection (wolf gets in) |
| Consequence | Adopting an ineffective treatment | Missing an effective treatment |
| Controlled by | Significance level (α), conventionally 0.05 | Statistical power (1 − β), conventionally 0.80 |
| Relationship to sample size | Fixed by α level | ↓ as sample size ↑ (more power) |
| Clinical impact | Patient receives unnecessary/ineffective treatment | Patient misses a beneficial treatment |
| Multiple comparisons | Risk increases with more tests (Bonferroni correction) | |
| Exam mnemonic | Type I = I made it up (false positive) | Type II = Too blind to see it (false negative) |
| Which is worse? | Context-dependent: worse in safety trials (approving a harmful drug) | Context-dependent: worse in screening (missing a treatable cancer) |
The trade-off: Reducing α (stricter threshold) decreases Type I error but increases Type II error (and vice versa). The only way to reduce BOTH is to increase sample size.
Table 3: Sensitivity vs Specificity vs PPV vs NPV
| Measure | Formula | Question Answered | Affected By |
|---|---|---|---|
| Sensitivity | TP / (TP + FN) | Of all DISEASED, how many tested positive? | Intrinsic test property |
| Specificity | TN / (FP + TN) | Of all HEALTHY, how many tested negative? | Intrinsic test property |
| PPV | TP / (TP + FP) | Of all TEST POSITIVES, how many truly have disease? | Prevalence ↑ → PPV ↑ |
| NPV | TN / (FN + TN) | Of all TEST NEGATIVES, how many are truly healthy? | Prevalence ↑ → NPV ↓ |
| Clinical Rule | Meaning | Use |
|---|---|---|
| SnNOut | Sensitive test + Negative result = rules OUT disease | Screening (don't want to miss cases) |
| SpPIn | Specific test + Positive result = rules IN disease | Confirmation (don't want false positives) |
| Measure | High Value Means | Clinical Application |
|---|---|---|
| LR+ (Sensitivity / (1-Specificity)) | Strong positive test performance | LR+ > 10 = very useful positive result |
| LR− ((1-Sensitivity) / Specificity) | Strong negative test performance | LR− < 0.1 = very useful negative result |
Critical insight: PPV and NPV depend on disease prevalence. Even a test with 99% sensitivity and 99% specificity will have poor PPV in a low-prevalence population because the absolute number of false positives will overwhelm true positives.
Table 4: Study Designs Comparison
| Feature | Cross-sectional | Case-Control | Cohort | RCT |
|---|---|---|---|---|
| Direction | Snapshot (no time direction) | Retrospective (outcome → exposure) | Prospective or retrospective (exposure → outcome) | Prospective (intervention → outcome) |
| What it measures | Prevalence | Odds Ratio (OR) | Relative Risk (RR), Incidence | Relative Risk, NNT, NNH |
| Starting point | Defined population | Cases + Controls selected | Exposed + Unexposed groups identified | Participants randomised |
| Can establish temporality | No | No (inferred) | Yes | Yes |
| Can establish causation | No | No (association only) | Stronger association | Yes (strongest) |
| Best for | Prevalence estimation, generating hypotheses | Rare diseases | Rare exposures, incidence | Testing interventions |
| Sample size needed | Moderate | Small–moderate | Large | Large |
| Time required | Short (one point) | Short–moderate | Long (years of follow-up) | Moderate–long |
| Cost | Low | Low–moderate | High | Very high |
| Key bias risk | Prevalence bias, temporal ambiguity | Recall bias, selection bias | Loss to follow-up, confounding | Hawthorne effect, strict inclusion → low external validity |
| Psychiatry example | NMHS 2016 prevalence survey | Trauma in BPD vs controls | Dunedin cohort study | STAR*D, CATIE |
| Level of evidence | Lower (descriptive) | Level 3b | Level 2b | Level 1b |
| Ethical concern | Minimal | Minimal | Minimal (observational) | Cannot randomise harmful exposures |
Table 5: Qualitative vs Quantitative Research
| Feature | Quantitative | Qualitative |
|---|---|---|
| Philosophy | Positivism (objective, measurable reality) | Constructivism (subjective, multiple realities) |
| Approach | Deductive (theory → hypothesis → data) | Inductive (data → patterns → theory) |
| Research question | "How much? How many? Is there a difference?" | "What is the experience? How do people make sense of it?" |
| Data | Numbers (scores, counts, measurements) | Words (transcripts, field notes, narratives) |
| Sample size | Large (powered by calculation) | Small (purposive, until data saturation) |
| Sampling | Probability (random, stratified, cluster) | Non-probability (purposive, snowball, theoretical) |
| Data collection | Questionnaires, rating scales, structured interviews | In-depth interviews, focus groups, observation |
| Analysis | Statistical (t-test, ANOVA, regression) | Thematic analysis, grounded theory, IPA, content analysis |
| Researcher role | Detached, objective observer | Embedded, reflexive participant |
| Rigour criteria | Validity + Reliability | Credibility + Transferability + Dependability + Confirmability |
| Output | Effect sizes, p-values, confidence intervals | Themes, narratives, theoretical frameworks |
| Generalisability | High (if representative sample) | Low (context-dependent; transferability instead) |
| Causation | Can establish (experimental designs) | Cannot establish (explores meaning, not causation) |
| Replication | High (standardised methods) | Low (unique contexts) |
| Strengths | Precision, generalisability, causal inference | Depth, context, lived experience, "why" |
| Limitations | Reductionism, misses meaning, decontextualises | Not generalisable, time-consuming, researcher bias |
| Psychiatry example | RCT of SSRIs for depression | Phenomenological study of voice-hearing experience |
| Mixed methods | Can combine both: quantitative outcomes + qualitative process evaluation (e.g., MANAS trial) |
Table 6: Systematic Review vs Meta-analysis vs Network Meta-analysis
| Feature | Systematic Review | Meta-analysis | Network Meta-analysis |
|---|---|---|---|
| Definition | Structured, reproducible synthesis of all evidence on a question | Statistical pooling of results from multiple studies into a single estimate | Extension of meta-analysis comparing multiple treatments simultaneously |
| Nature of synthesis | Qualitative (narrative) | Quantitative (statistical) | Quantitative (statistical, complex) |
| Includes statistical pooling | Not necessarily | Yes (always) | Yes (always) |
| Types of evidence compared | All studies on one question | Direct comparisons (A vs B) | Direct AND indirect comparisons (A vs B, B vs C → infer A vs C) |
| Key output | Summary tables, quality assessment, narrative conclusion | Pooled effect size, forest plot, I2 | Treatment rankings, league table, network diagram |
| Heterogeneity assessment | Described qualitatively | I2, Cochrane Q | I2, inconsistency, node-splitting |
| Publication bias | Discussed qualitatively | Funnel plot, Egger's test | Comparison-adjusted funnel plot |
| When NOT to pool | When heterogeneity is extreme | When transitivity assumption is violated | |
| Key assumption | Comprehensive search, reproducible methods | Studies are sufficiently similar to pool | Transitivity (similar populations across comparisons) |
| Reporting guideline | PRISMA | PRISMA | PRISMA-NMA |
| Psychiatry example | Cochrane review of CBT for depression | Cipriani et al., 2018 (antidepressants) | Cipriani et al., 2018 (21 antidepressants ranked) |
| Level of evidence | 1a | 1a | 1a |
Table 7: Probability vs Non-Probability Sampling
| Feature | Probability Sampling | Non-Probability Sampling |
|---|---|---|
| Definition | Every member of the population has a known, non-zero chance of selection | Selection is based on availability, judgment, or other non-random criteria |
| Basis | Random selection | Convenience, judgment, or self-selection |
| Generalisability | High (representative of population) | Low (may not represent population) |
| Bias risk | Lower (randomisation minimises selection bias) | Higher (selection bias inherent) |
| Sampling frame | Required (complete list of population) | Not required |
| Cost and effort | Higher (need complete frame, randomisation) | Lower (practical, quick) |
| Use in research | Quantitative studies, surveys, clinical trials | Qualitative studies, pilot studies, hard-to-reach populations |
| Type | Method | Example |
|---|---|---|
| Simple random | Every individual has equal chance | Random number table to select patients from hospital register |
| Stratified | Divide into strata, random sample from each | Stratify by age group, randomly sample within each stratum |
| Cluster | Randomly select entire clusters | Randomly select 10 PHCs from a district, study all patients at those PHCs |
| Systematic | Every kth individual | Every 5th patient from the OPD register |
| Convenience | Whoever is available | Patients attending OPD on the day of data collection |
| Purposive | Researcher selects based on characteristics | Selecting only patients with treatment-resistant depression for qualitative interviews |
| Snowball | Participants recruit others | Studying IV drug users, each participant refers peers |
| Quota | Non-random but ensures proportions | Ensuring 50% male and 50% female in sample (non-randomly selected within each group) |
Table 8: Types of Bias in Research
| Bias Type | Definition | Stage | Example | Prevention |
|---|---|---|---|---|
| Selection bias | Systematic difference between study participants and target population | Design | Only enrolling tertiary hospital patients (sicker, more complex) | Random sampling, clear inclusion criteria |
| Recall bias | Cases recall exposure differently from controls | Data collection | Mothers of children with defects remember medication use more accurately | Prospective design, objective records |
| Observer/Detection bias | Assessor's expectations influence measurement | Data collection | Unblinded rater scores drug group as improved | Blinding of assessors |
| Performance bias | Differential treatment between groups (beyond the intervention) | Conduct | Intervention group receives more clinician attention | Double blinding, standardised protocols |
| Attrition bias | Differential loss to follow-up between groups | Follow-up | Sicker patients drop out of drug arm | ITT analysis, minimise dropout |
| Publication bias | Studies with positive results are more likely published | Reporting | Only significant antidepressant trials in literature | Trial registries, grey lit search, funnel plot |
| Confounding | Third variable associated with both exposure and outcome | Analysis | Coffee → lung cancer (confounder: smoking) | Randomisation, stratification, multivariate analysis |
| Lead-time bias | Earlier detection inflates apparent survival time | Interpretation | Screening detects cancer 2 years earlier without changing mortality | Measure mortality, not survival from diagnosis |
| Berkson bias | Spurious association in hospital-based studies | Design | Depression-diabetes association inflated because both increase hospitalisation | Population-based studies |
| Information bias | Systematic errors in data collection | Data collection | Using an unvalidated questionnaire | Standardised, validated instruments |
| Reporting bias | Selective reporting of outcomes | Reporting | Reporting only the outcome that was significant | Pre-registration, protocol publication |
| Hawthorne effect | Participants change behaviour because they are observed | Conduct | Patients adhere better during a trial than in routine care | Naturalistic designs, long follow-up |
| Volunteer bias | Volunteers differ systematically from non-volunteers | Design | Volunteers are healthier, more motivated, more educated | Random sampling (when possible) |
Cross-reference: D1 (Study Notes), D2 (Model Answers), D3 (Mnemonics), D6 (Quick Review)
PYQ Frequency Analysis
Source: PG exams Dec 2011, Jun 2025 + PG exams 2013-2022
Executive Summary
Statistics and research methodology is a consistently tested cluster, arguably the most predictable in Paper I after neurotransmitters and sleep. The questions follow a very stable pattern: validity/reliability, study designs, statistical tests, and evidence-based medicine rotate reliably.
Key insight: Unlike clinical topics, statistics questions have almost no variability. The same 5-6 question templates recycle. Master those templates and you're covered.
Topic-Level Frequency
| Topic | Exam Mentions | Avg per Exam | Verdict |
|---|---|---|---|
| Validity & Reliability | 5 | ~0.18 | Every 5-6 exams |
| Epidemiological study designs | 4 | ~0.14 | Every 6-7 exams |
| Statistical tests (t-test, Chi-square, ANOVA, non-parametric) | 5 | ~0.18 | Every 5-6 exams |
| Evidence-based medicine | 5 | ~0.18 | Every 5-6 exams |
| Qualitative vs Quantitative research | 3 | ~0.11 | Every 8-9 exams |
| Meta-analysis / Systematic review | 3 | ~0.11 | Emerging topic |
| Sampling & Sample size | 2 | ~0.07 | Occasional |
| Combined cluster | ~27 | ~0.96 | ~1 question per exam |
Key PYQs Identified
Validity & Reliability
- "Define validity. Name different types. Importance in psychiatry." [10 marks, split 2+6+2]
- "What is validity? Enumerate types. Criterion validity." [10 marks, split 2+3+5]
- "What is validity? Discuss different types with examples." [10 marks, split 3+7]
- "Define reliability and validity. Significance. Types." [10 marks, split 3+2+5]
- "What is reliability? Types with examples. Internal and external validity." [10 marks, split 2+4+4]
Study Designs & Epidemiology
- "What is psychiatric epidemiology? Uses. Types of designs." [10 marks, split 2+4+4]
- "Describe different types of epidemiological studies in Psychiatry." [10 marks]
- "What is genetic epidemiology? Three types of genetic studies." [10 marks]
- "Define epidemiology. Different types of epidemiological studies." [10 marks]
Statistical Tests
- "Discuss Non-Parametric Tests." [10 marks]
- "Chi Square test." [10 marks]
- "Paired t Test." [10 marks]
- "Enumerate statistical tests to measure differences. Describe t-test with example." [10 marks, split 4+6]
- "Describe Factor Analysis and its applications in psychiatric research." [10 marks]
- "Enumerate statistical tests. Describe paired t-test. Two non-parametric tests for group difference." [10 marks, split 4+2+4]
Evidence-Based Medicine
- "What is evidence based medicine? How is it practiced? Steps with example." [10 marks, split 2+2+6]
- "What is EBM? How is it practiced in India? Limitations." [10 marks]
- "Discuss levels of evidence and their role in EBM." [10 marks]
- "What is EBM? What is personalized medicine? Pharmacogenetics." [10 marks, split 2+3+5]
Research Methodology
- "What is quantitative research? Essential features, strengths, limitations." [10 marks]
- "What is qualitative research? How does it differ from quantitative?" [10 marks, split 2+5+3]
- "What is meta-analysis? Systematic review? Key meta-analytical studies in psychiatry." [10 marks]
- "Odds ratio. Confounding, effect modification, mediation." [10 marks, split 4+2+2+2]
- "How to generate research question? Types of hypothesis. Meta-analysis vs network analysis." [10 marks]
- "Sampling techniques. Sample size. Principles of sample size calculation." [10 marks, split 2+2+6]
Long Essay Candidates
| Rank | Topic | Probability |
|---|---|---|
| 1 | "Define validity and reliability. Types. Clinical importance." | Very High |
| 2 | "What is EBM? Steps. Practice in India. Limitations." | Very High |
| 3 | "Enumerate statistical tests. Describe t-test/chi-square with examples." | High |
| 4 | "Describe different epidemiological study designs in psychiatry." | High |
| 5 | "Qualitative vs quantitative research, features, strengths, limitations." | Medium-High |
Exam Strategy
Must-Prepare (will be asked)
- Validity, define, 4 types (face, content, criterion [concurrent + predictive], construct [convergent + discriminant]), examples from psychiatry
- Reliability, define, types (test-retest, inter-rater, split-half, internal consistency/Cronbach's alpha), relationship to validity
- Statistical tests, when to use each: t-test (2 means, parametric), ANOVA (>2 groups), Chi-square (categorical data), Mann-Whitney/Wilcoxon (non-parametric equivalents)
- EBM, define, 5 steps (ask, acquire, appraise, apply, assess), levels of evidence pyramid
Should-Prepare
- Study designs, cross-sectional, case-control, cohort, RCT, ecological. Strengths/limitations table
- Meta-analysis, define, steps, forest plot, heterogeneity (I2), advantages over narrative review
- Sensitivity/Specificity/PPV/NPV, 2x2 table, formulae, clinical application
Nice-to-Know
- Factor analysis (exploratory vs confirmatory)
- Odds ratio vs relative risk
- Confounding and effect modification
- Sampling techniques (probability vs non-probability)
Emerging Trends (2020+)
- Network meta-analysis, appeared Jun 2025
- Pharmacogenetics + personalized medicine, combining EBM with genetics
- Yoga in psychiatry + methodological issues, India-specific
- Sample size calculation, increasingly tested
Analysis based on PG exams Dec 2011, Jun 2025 + PG exams 2013-2022.
Quick Review
RECALL Questions (1–12)
Q1. Define face validity and give one limitation.
A: Face validity is the degree to which a test appears, on the surface, to measure what it intends to measure. It is assessed subjectively by non-experts. Limitation: It is the weakest form of validity, a test can look right but actually measure a different construct (e.g., a test that appears to measure anxiety might actually be measuring general distress).
Q2. What are the two subtypes of criterion validity?
A: (1) Concurrent validity, the test and the gold standard criterion are measured at the same time (e.g., a new depression scale administered alongside the HAM-D). (2) Predictive validity, the test predicts a future outcome (e.g., AUDIT score at admission predicting withdrawal severity during hospitalisation).
Q3. Name the four types of reliability and their associated statistics.
A:
- Test-retest → Pearson r or ICC
- Inter-rater → Cohen's kappa (κ)
- Split-half → Spearman-Brown corrected r
- Internal consistency → Cronbach's alpha (α)
Q4. What is the non-parametric alternative to the independent t-test?
A: The Mann-Whitney U test. It compares the rank distributions of two independent groups and does not assume normal distribution.
Q5. What is the non-parametric alternative to the paired t-test?
A: The Wilcoxon signed-rank test. It compares two related measurements by analysing the magnitude and direction of differences between pairs.
Q6. State the formula for sensitivity and specificity.
A: From a 2×2 table (a = TP, b = FP, c = FN, d = TN):
- Sensitivity = a / (a + c), proportion of true positives among all diseased
- Specificity = d / (b + d), proportion of true negatives among all non-diseased
Q7. What are the 5 steps of EBM?
A: Ask (formulate PICO question), Acquire (search for evidence), Appraise (critically evaluate), Apply (integrate with expertise and patient values), Assess (evaluate outcomes). The "5 A's."
Q8. What is the difference between a systematic review and a meta-analysis?
A: A systematic review is a structured, reproducible synthesis of all available evidence on a research question (qualitative/narrative). A meta-analysis is the statistical pooling of results from multiple studies into a single quantitative estimate. Every meta-analysis is part of a systematic review, but not every systematic review includes a meta-analysis (e.g., when studies are too heterogeneous to pool).
Q9. Define NNT and give its formula.
A: NNT (Number Needed to Treat) = the number of patients who need to be treated with the intervention for one additional patient to benefit compared to the control. Formula: NNT = 1 / ARR, where ARR (Absolute Risk Reduction) = Control Event Rate − Experimental Event Rate. Lower NNT = more effective treatment.
Q10. What is the PICO framework?
A: A structured format for clinical research questions: P = Population/Patient, I = Intervention, C = Comparison/Control, O = Outcome. Example: "In adults with MDD (P), does CBT + SSRI (I) compared to SSRI alone (C) improve remission rates (O)?"
Q11. Name three probability and three non-probability sampling methods.
A: Probability: Simple random, stratified random, cluster sampling. Non-probability: Convenience, purposive, snowball sampling.
Q12. What does an I2 of 75% mean in a meta-analysis?
A: I2 = 75% indicates substantial heterogeneity, 75% of the variability across study results is due to real differences between studies (heterogeneity) rather than chance. This suggests a random-effects model should be used, and sources of heterogeneity should be explored through subgroup analysis or meta-regression. (Interpretation: 0–25% low, 25–50% moderate, 50–75% substantial, >75% considerable.)
APPLICATION Questions (13–25)
Q13. A study finds p = 0.03 with a 95% CI of 1.2–3.4 for the odds ratio. Interpret.
A: The result is statistically significant (p < 0.05). The odds of the outcome are 1.2 to 3.4 times higher in the exposed group compared to the unexposed group (95% confidence). Since the CI does not include 1 (the null value for OR), this confirms statistical significance. The point estimate of the OR lies at some value between 1.2 and 3.4 (likely around 2.0). The exposure is a significant risk factor for the outcome.
Q14. A researcher wants to compare anxiety scores (normally distributed) before and after a 12-week yoga intervention in the same 40 patients. Which test?
A: Paired t-test. Rationale: The data is continuous, normally distributed, and involves two measurements from the same subjects (before-after design). If normality of differences were violated, use Wilcoxon signed-rank test instead.
Q15. A study compares treatment response (responder/non-responder) across three drug groups. Which test?
A: Chi-square test of independence. Rationale: The outcome is categorical (responder vs non-responder) and there are three independent groups. This creates a 3×2 contingency table. If any expected cell count is < 5, use Fisher's exact test or collapse categories.
Q16. You want to compare depression scores (skewed distribution) across four therapy groups with 12 patients each. Which test?
A: Kruskal-Wallis test. Rationale: Four independent groups, continuous but non-normally distributed data, relatively small sample sizes. This is the non-parametric alternative to one-way ANOVA. If significant, follow up with Dunn's test for pairwise comparisons.
Q17. A case-control study shows OR = 4.5 (crude) for the association between childhood trauma and BPD. After adjusting for parental substance use, the OR becomes 2.1. What happened?
A: Confounding was present. Parental substance use was a confounder, it was associated with both the exposure (childhood trauma) and the outcome (BPD). The crude OR of 4.5 was inflated by the confounding effect. The adjusted OR of 2.1 represents the true association after removing the confounding effect of parental substance use. Since the crude and adjusted estimates differ substantially (> 10% change), confounding is confirmed.
Q18. A screening test has sensitivity = 95% and specificity = 90%. In a population with 1% prevalence of the disease, calculate the approximate PPV.
A: For 10,000 people: 100 have disease, 9,900 do not.
- TP = 95% × 100 = 95
- FP = 10% × 9,900 = 990
- PPV = 95 / (95 + 990) = 95 / 1,085 = 8.8%
Despite excellent sensitivity and specificity, the PPV is only 8.8% because the disease is rare. Most positive results are false positives. This is why screening for low-prevalence conditions generates many false alarms.
Q19. In an RCT, 60% of the control group relapsed and 40% of the drug group relapsed. Calculate the NNT.
A:
- ARR = CER − EER = 0.60 − 0.40 = 0.20
- NNT = 1 / 0.20 = 5
Interpretation: You need to treat 5 patients with the drug to prevent one additional relapse compared to control. This is a clinically meaningful NNT.
Q20. A forest plot shows 8 studies. 5 studies have confidence intervals that cross the line of no effect. The summary diamond does NOT cross the line of no effect. Is the overall result significant?
A: Yes, the overall result is statistically significant. While individual studies may be non-significant (their CIs cross the null line), the pooled estimate (summary diamond) does not cross the line of no effect, indicating that when all evidence is combined, the treatment effect is statistically significant. This demonstrates the power of meta-analysis, pooling data from multiple underpowered individual studies can reveal a significant effect.
Q21. A researcher measures the correlation between BDI scores and hours of sleep per night in 200 patients with depression. BDI scores are normally distributed. Which test?
A: Pearson's correlation coefficient (r). Rationale: Both variables are continuous (BDI score and hours of sleep), data is normally distributed, and the researcher is examining a linear association between two variables. If either variable were non-normally distributed or ordinal, use Spearman's ρ instead.
Q22. A study of an SSRI vs placebo reports: "Treatment group improved by 3 points more on HAM-D (p = 0.04, Cohen's d = 0.15)." Is this clinically meaningful?
A: Likely not clinically meaningful despite statistical significance. While p = 0.04 is statistically significant, Cohen's d = 0.15 indicates a very small effect size (small = 0.2, medium = 0.5, large = 0.8). The 3-point HAM-D difference is at the borderline of the minimum clinically important difference (MCID, typically 3–4 points). This illustrates that statistical significance does not equal clinical significance, with a large enough sample, even trivially small differences become statistically significant. Clinical decision-making should weigh effect size and MCID, not just p-values.
Q23. A researcher wants to study the experience of living with treatment-resistant depression from the patient's perspective. Which research methodology is most appropriate?
A: Qualitative phenomenological study using in-depth semi-structured interviews. Phenomenology focuses on the lived experience of a phenomenon and aims to understand its essential meaning from the participant's perspective. Data would be analysed using interpretative phenomenological analysis (IPA) or Colaizzi's method. The sample would be small (8–15 participants), purposively selected, and data collection would continue until thematic saturation.
Q24. An RCT randomises 200 patients to drug vs placebo. During the trial, 30 patients in the drug arm stop taking the medication due to side effects. How should they be analysed?
A: Under Intention-to-Treat (ITT) analysis, all 200 patients are analysed in their originally assigned groups, the 30 who discontinued are still counted in the drug arm. This preserves the benefits of randomisation and provides a conservative estimate of treatment effect. ITT reflects real-world effectiveness (not all patients comply). A per-protocol analysis (excluding non-compliers) can be reported as secondary analysis but may overestimate the true treatment effect by excluding those who had side effects.
Q25. A funnel plot in a meta-analysis shows asymmetry, with missing studies in the lower-left region. What does this suggest?
A: This suggests publication bias, smaller studies with negative or null results (lower-left region = small sample size + small/negative effect) are missing from the literature, likely because they were never published. This means the pooled estimate may overestimate the true treatment effect. The researcher should: (1) run Egger's test to confirm statistically, (2) apply trim-and-fill method to adjust the estimate, (3) search for grey literature and unpublished data.
ANALYSIS Questions (26–35)
Q26. Why might a highly sensitive test have poor PPV in a low-prevalence population?
A: Because PPV depends on both the test's specificity and the disease prevalence. In a low-prevalence population, the vast majority of people are disease-free. Even with high sensitivity (catching nearly all true cases), the small number of true positives is overwhelmed by false positives from the large healthy population. For example, if sensitivity = 99% and specificity = 95%, and prevalence = 0.1%: in 100,000 people, 100 have disease (99 detected) but 4,995 healthy people test falsely positive. PPV = 99/(99+4,995) = 1.9%. The absolute number of false positives is determined by (1-specificity) × (number of healthy people), which is enormous when prevalence is low.
Q27. A researcher reports that Cohen's kappa for their diagnostic interview is 0.45. Is this adequate for clinical use? Why or why not?
A: κ = 0.45 indicates moderate agreement, which is generally not adequate for clinical use, especially for high-stakes decisions. In clinical diagnostics, substantial to almost perfect agreement (κ > 0.60, ideally > 0.80) is expected. A κ of 0.45 means raters disagree on a substantial proportion of cases even after accounting for chance agreement. This could lead to misdiagnosis depending on which clinician the patient sees. The instrument or training protocol needs improvement before clinical deployment. However, for screening purposes (lower stakes, further assessment follows), moderate reliability may be acceptable.
Q28. An RCT of a new antidepressant excludes patients with comorbid substance use, personality disorders, and suicidal ideation. How does this affect the study's validity?
A: This creates a tension between internal and external validity:
- Internal validity increases: By excluding complex patients, the study reduces confounders (substance use affecting response, personality pathology affecting adherence, suicidality requiring additional interventions). This creates a cleaner test of the drug's efficacy.
- External validity decreases dramatically: In real-world psychiatry, comorbidity is the rule, not the exception. Approximately 40–60% of patients with depression have comorbid substance use, personality difficulties, or suicidal ideation. The study sample represents a highly selected, atypical population. Results cannot be confidently generalised to the patients most commonly seen in clinical practice. This is the "efficacy-effectiveness gap", drugs that work in clinical trials may not work as well in routine care.
Q29. A researcher finds a statistically significant correlation (r = 0.15, p = 0.01) between social media use and depression scores in a sample of 5,000 adolescents. What should you conclude?
A: The finding is statistically significant but clinically trivial. r = 0.15 is a very weak correlation, it explains only 2.25% of the variance in depression scores (r2 = 0.0225). The statistical significance is driven entirely by the large sample size (n = 5,000), which gives power to detect even negligible effects. The 97.75% of variance is explained by other factors. Additionally, this is a correlation, not causation, it could be that (a) social media causes depression, (b) depression leads to more social media use, or (c) a third variable (e.g., loneliness, sleep deprivation) drives both. This highlights why effect size and clinical significance matter more than p-values.
Q30. Why is ITT analysis considered more conservative than per-protocol analysis in a superiority trial?
A: ITT analysis includes ALL randomised participants in their assigned groups, regardless of whether they completed the intervention or adhered to the protocol. This is conservative because:
- Non-compliers dilute the treatment effect, patients who stopped the drug (due to side effects, lack of response, etc.) are still counted in the drug group, pulling the drug group's outcome toward the control group's outcome.
- Preserves randomisation, maintaining the original randomised groups ensures that known and unknown confounders remain balanced.
- Reflects real-world effectiveness, in practice, not all patients comply, so ITT gives a more realistic estimate.
Per-protocol analysis, by excluding non-compliers, creates a biased sample (those who tolerated and adhered may be inherently different) and can overestimate the treatment effect. However, in non-inferiority trials, per-protocol is actually the primary analysis because ITT (by diluting differences) can falsely support non-inferiority.
Q31. A new psychiatric rating scale has a Cronbach's alpha of 0.97. Is this a problem?
A: Potentially, yes. While α > 0.70 indicates acceptable internal consistency, an α of 0.97 suggests item redundancy, the items are so highly intercorrelated that they are essentially asking the same thing in slightly different ways. This means the scale is longer than necessary without adding new information. It increases respondent burden, completion time, and the risk of response fatigue without improving measurement. The scale developer should consider reducing items through item-total correlation analysis, removing items with very high inter-item correlations (> 0.90) that are redundant. An optimal Cronbach's alpha is typically 0.80–0.90.
Q32. A case-control study finds that patients with schizophrenia are more likely to have been born in winter months. What type of bias could explain this finding, and what design would strengthen the evidence?
A: This could be a genuine finding (season of birth effect is well-replicated) but could also be affected by Berkson bias (if controls were hospital-based with different seasonal admission patterns) or selection bias (if the sampling frame introduced seasonal artefacts). Recall bias is less relevant here since date of birth is an objective fact from records. To strengthen the evidence: (1) A large prospective birth cohort study following individuals from birth would establish temporality without selection bias. (2) Using population-based controls rather than hospital controls eliminates Berkson bias. (3) The ecological fallacy should be considered if aggregate population-level birth data is used. The Scandinavian birth register studies (e.g., Danish cohort) provide the strongest evidence for this association.
Q33. Why is the random-effects model preferred over the fixed-effect model when heterogeneity is high in a meta-analysis?
A: The fixed-effect model assumes there is ONE true effect size and all variation between studies is due to sampling error (chance). The random-effects model assumes the true effect varies between studies (because of differences in populations, interventions, settings) and accounts for both within-study and between-study variance.
When heterogeneity is high (I2 > 50%), the fixed-effect assumption is violated, studies clearly differ in their true effects. Using a fixed-effect model would produce inappropriately narrow confidence intervals (false precision) because it ignores the real variability between studies. The random-effects model produces wider, more honest confidence intervals that reflect the genuine uncertainty. It also weights studies more equally (small studies get relatively more weight compared to fixed-effect), which can be both an advantage (less dominated by large studies) and a limitation (gives more weight to potentially lower-quality small studies).
Q34. A researcher performs 20 independent t-tests on the same dataset comparing treatment vs control on 20 different outcome measures. All are tested at α = 0.05. What is the problem and what is the solution?
A: The problem is multiple comparisons (inflated Type I error). With 20 independent tests at α = 0.05, the probability of at least one false positive is: 1 − (1 − 0.05)20 = 1 − 0.36 = 0.64 (64%). There is a 64% chance of finding at least one "significant" result by chance alone, even if there is no true effect.
Solutions:
- Bonferroni correction: Adjust α to 0.05/20 = 0.0025 per test. Simple but very conservative, increases Type II error.
- Holm-Bonferroni (step-down): Less conservative, ranks p-values and applies progressively less strict thresholds.
- False Discovery Rate (FDR) control (Benjamini-Hochberg): Controls the proportion of false positives among significant results rather than the overall error rate. Less conservative than Bonferroni.
- Pre-specify a primary outcome: Designate one primary outcome and treat others as secondary/exploratory, avoids the multiple comparison problem for the primary endpoint.
- MANOVA: Use multivariate ANOVA to test all outcomes simultaneously.
Q35. Explain why the STAR*D trial is considered methodologically important in psychiatric research, linking it to concepts of internal and external validity.
A: The STAR*D (Sequenced Treatment Alternatives to Relieve Depression) trial is landmark because it prioritised external validity in a field dominated by explanatory RCTs with strict inclusion criteria.
External validity strengths:
- Enrolled patients from both primary care and psychiatric settings (real-world recruitment)
- Minimal exclusion criteria, included patients with comorbid anxiety, substance use, and medical conditions (unlike most antidepressant trials)
- Sequential treatment design reflecting real clinical decision-making (what to do when first treatment fails)
- Measured remission (HAM-D ≤ 7) as primary outcome rather than just response (more clinically meaningful)
Internal validity trade-offs:
- No placebo control in most steps (ethical and practical decision, but limits causal inference about specific treatments)
- Not blinded in all steps (participants chose some treatment options)
- High dropout rate (28% completed all steps), potential attrition bias
Key findings that shaped clinical practice:
- Overall remission rate across all steps was approximately 67% (cumulative)
- Remission rates decreased with each subsequent step (37% → 31% → 14% → 13%)
- Demonstrated that a substantial proportion of patients need treatment modification
- Provided the first large-scale sequential algorithm data for treatment-resistant depression
STAR*D exemplifies the efficacy vs effectiveness distinction: it sacrificed some internal validity (no placebo arm, open-label steps) to maximise external validity (generalisable to real patients). This makes its findings more directly applicable to clinical practice than most tightly controlled explanatory RCTs.
Cross-reference: D1 (Study Notes), D2 (Model Answers), D3 (Mnemonics), D4 (Comparisons)