← All study guides
Guide 04 · Part I

Research Methods

Paper I · Basic Sciences. Six study modes, from notes to quick review.

Most askedvalidity and reliability typesevidence-based medicine stepsstatistical tests selectionepidemiological study designsqualitative vs quantitative researchmeta-analysis systematic review
Jump to a section
Chapter 01

Study Notes



1. VALIDITY

1.1 Definition

Validity = the degree to which a test or instrument measures what it claims to measure. A valid IQ test actually measures intelligence, not reading speed.

1.2 Types of Validity

TypeDefinitionExample in Psychiatry
FaceAppears to measure what it claims (subjective judgment)PHQ-9 "looks like" a depression questionnaire
ContentCovers the full domain of the constructHAM-D covers mood, guilt, insomnia, psychomotor, somatic, the full depression domain
CriterionCorrelates with an external gold standardMMSE score correlates with clinical dementia diagnosis
ConstructMeasures the theoretical construct accuratelyBDI correlates with other depression measures (convergent) but not with anxiety measures (discriminant)

1.3 Criterion Validity: Two Subtypes

SubtypeTimingExample
ConcurrentTest and criterion measured at the same timeNew depression scale given alongside HAM-D to same patients
PredictiveTest predicts a future outcomeGRE score predicting future academic performance

1.4 Construct Validity: Two Subtypes

SubtypeDefinitionExample
ConvergentCorrelates with measures of the same constructBDI correlates highly with PHQ-9 (both measure depression)
DiscriminantDoes NOT correlate with measures of different constructsBDI should not correlate highly with a mania scale

1.5 Internal vs External Validity

FeatureInternal ValidityExternal Validity
DefinitionDegree to which results are due to the independent variable, not confoundersDegree to which results can be generalised to other populations/settings
FocusCause-and-effect accuracyApplicability beyond the study
Stronger inRCTs (controlled environment)Community-based studies, pragmatic trials
ThreatsSelection bias, maturation, testing effects, history, attrition, instrumentation, regression to meanSampling bias, Hawthorne effect, setting specificity, volunteer bias
Trade-offTight control increases internal but may reduce externalBroad inclusion increases external but may reduce internal

1.6 Threats to Internal Validity (Mnemonic: SHRIMP-T)

1.7 Threats to External Validity


2. RELIABILITY

2.1 Definition

Reliability = the degree to which a test produces consistent, reproducible results across time, raters, or items. A reliable thermometer gives the same reading each time you measure the same temperature.

2.2 Types of Reliability

TypeWhat It MeasuresStatisticAcceptable Value
Test-retestStability over time (same test, same people, different times)Pearson r or ICCr > 0.7
Inter-raterAgreement between two or more ratersCohen's kappa (κ)κ > 0.6
Split-halfConsistency between two halves of the same testSpearman-Brown corrected rr > 0.7
Internal consistencyHow well items within a test measure the same constructCronbach's alpha (α)α > 0.7

2.3 Cohen's Kappa: Interpretation

Kappa Value · Interpretation
< 0.00 Less than chance agreement
0.01–0.20 Slight agreement
0.21–0.40 Fair agreement
0.41–0.60 Moderate agreement
0.61–0.80 Substantial agreement
0.81–1.00 Almost perfect agreement

Why kappa over simple percentage agreement? Kappa accounts for agreement expected by chance alone. Two raters flipping coins would agree 50% of the time, kappa corrects for this.

2.4 Cronbach's Alpha: Key Points

2.5 Reliability vs Validity: The Relationship

Statement · True/False
A test can be reliable but not valid TRUE, a scale can consistently measure the wrong thing
A test can be valid but not reliable FALSE, if it measures the right thing, it must do so consistently
Reliability is necessary but not sufficient for validity TRUE
High reliability guarantees high validity FALSE

Classic analogy: A marksman shooting a tight cluster (reliable) but missing the bullseye (not valid). Scattered shots hitting the bullseye on average are neither reliable nor truly valid for individual measurement.


3. STUDY DESIGNS

3.1 Hierarchy of Evidence (strongest to weakest)

  1. Systematic review / Meta-analysis
  2. Randomised Controlled Trial (RCT)
  3. Cohort study
  4. Case-control study
  5. Cross-sectional study
  6. Ecological study
  7. Case series / Case report
  8. Expert opinion

3.2 Study Design Comparison Table

FeatureCross-sectionalCase-controlCohortRCT
DirectionSnapshotBackward (retrospective)Forward (prospective or retrospective)Forward (prospective)
MeasurePrevalenceOdds ratioRelative risk, incidenceRelative risk, NNT
Exposure & outcomeMeasured simultaneouslyOutcome known, exposure soughtExposure known, outcome awaitedExposure assigned, outcome awaited
CausationCannot establishSuggests associationStronger associationStrongest causation
TimeOne pointPast exposure recalledLong follow-upFixed duration
CostLowModerateHighVery high
Best forPrevalence estimation, generating hypothesesRare diseasesRare exposures, incidence estimationTesting interventions

3.3 Cross-sectional Study

3.4 Case-Control Study

Interpreting Odds Ratio:

3.5 Cohort Study

Interpreting Relative Risk:

RR vs OR: In rare diseases (prevalence < 10%), OR approximates RR. In common diseases, OR overestimates RR.

3.6 Randomised Controlled Trial (RCT)

Key Components:

Component · Purpose
Randomisation Eliminates selection bias, balances known and unknown confounders
Allocation concealment Prevents foreknowledge of assignment (different from blinding)
Single blinding Participant blinded
Double blinding Participant + assessor blinded
Triple blinding Participant + assessor + statistician blinded
ITT (Intention-to-treat) Analyse all participants in their assigned group regardless of compliance, preserves randomisation, conservative estimate
Per-protocol Analyse only those who completed the protocol, can overestimate treatment effect

3.7 Ecological Study

3.8 Case Series and Case Report

FeatureCase ReportCase Series
N1 patientMultiple patients (no control group)
UseNovel presentations, rare side effects, new treatment ideasDescribe patterns, generate hypotheses
StrengthFirst signal of new phenomenonIdentifies emerging patterns
LimitationNo comparison, no generalisabilityNo controls, no causation
ExampleFirst report of NMS with a new antipsychoticSeries of 20 patients with treatment-resistant depression who responded to ketamine

4. STATISTICAL TESTS

4.1 Key Concepts First

Term · Definition
p-value Probability of obtaining results as extreme as observed, assuming H0 is true. p < 0.05 = statistically significant
Confidence Interval (CI) Range within which the true population parameter lies with 95% probability. If 95% CI for OR does not include 1 significant
Effect size Magnitude of the difference (Cohen's d, η2, r). Small = 0.2, Medium = 0.5, Large = 0.8
Power Probability of detecting a true effect (1 − β). Aim for ≥ 0.80
Type I error (α) Rejecting H0 when it is true (false positive). Threshold typically 0.05
Type II error (β) Failing to reject H0 when it is false (false negative). Related to power

4.2 The Decision Tree: Which Test When?

Step 1: What type of data?

Data TypeExamplesTests
Continuous (interval/ratio)HAM-D scores, age, weightt-test, ANOVA, Pearson r
Categorical (nominal/ordinal)Diagnosis (yes/no), severity (mild/moderate/severe)Chi-square, Fisher's exact

Step 2: How many groups?

GroupsContinuous (Parametric)Continuous (Non-parametric)Categorical
2 groups, unpairedIndependent t-testMann-Whitney UChi-square
2 groups, pairedPaired t-testWilcoxon signed-rankMcNemar's test
3+ groups, unpairedOne-way ANOVAKruskal-WallisChi-square
3+ groups, pairedRepeated measures ANOVAFriedman testCochran's Q

Step 3: Parametric or non-parametric?

Use parametric tests when:

Use non-parametric tests when:

4.3 Individual Tests: Details

t-test (Student's t-test)

Independent (unpaired) t-test:

Paired t-test:

ANOVA (Analysis of Variance)

One-way ANOVA:

Two-way ANOVA:

Repeated measures ANOVA:

Chi-Square Test

Test of Independence:

Goodness of Fit:

If expected cell count < 5 use Fisher's exact test

Non-Parametric Tests

Mann-Whitney U test:

Wilcoxon signed-rank test:

Kruskal-Wallis test:

Correlation
FeaturePearson rSpearman ρ
Data typeContinuous, normally distributedOrdinal or non-normal continuous
MeasuresLinear associationMonotonic association
Range−1 to +1−1 to +1
ExampleCorrelation between HAM-D and BDI scoresCorrelation between severity ranking and treatment response ranking

Interpreting r:

Correlation = Causation. Two variables may correlate because of a shared confounder.

Fisher's Exact Test
Regression
TypeUseExample
Linear regressionPredict continuous outcome from one or more predictorsPredicting HAM-D score from age, duration of illness, baseline score
Logistic regressionPredict binary outcome from one or more predictorsPredicting treatment response (yes/no) from baseline variables
Multiple regressionMultiple predictors for one outcomeEffect of CBT hours, medication adherence, and social support on depression score

5. SENSITIVITY AND SPECIFICITY

5.1 The 2×2 Table

Disease +Disease −
Test +a (True Positive)b (False Positive)a + b
Test −c (False Negative)d (True Negative)c + d
a + cb + dN

5.2 Formulae

MeasureFormulaMeaning
Sensitivitya / (a + c)Proportion of actual positives correctly identified (true positive rate)
Specificityd / (b + d)Proportion of actual negatives correctly identified (true negative rate)
PPVa / (a + b)Probability that a positive test = true disease
NPVd / (c + d)Probability that a negative test = true absence
LR+Sensitivity / (1 − Specificity)How much a positive test increases the odds of disease
LR−(1 − Sensitivity) / SpecificityHow much a negative test decreases the odds of disease

5.3 Clinical Rules

Why? A highly sensitive test catches almost all true cases very few false negatives if it's negative, you can be confident the disease is absent.

5.4 Effect of Prevalence on PPV and NPV

PrevalencePPVNPV
High (common disease)↑ Higher↓ Lower
Low (rare disease)↓ Lower↑ Higher

Clinical implication: Even a highly sensitive and specific test will have poor PPV in a low-prevalence population because the absolute number of false positives exceeds true positives.

Example: If PHQ-9 sensitivity = 90%, specificity = 85%, and depression prevalence = 2%:

5.5 ROC Curve


6. EVIDENCE-BASED MEDICINE (EBM)

6.1 Definition

EBM = the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients. Integrates:

  1. Best available evidence
  2. Clinical expertise
  3. Patient values and preferences

(David Sackett, 1996)

6.2 Five Steps of EBM (Mnemonic: 5 A's)

StepActionExample
1. AskFormulate a PICO questionIn adults with MDD (P), does CBT + SSRI (I) compared to SSRI alone (C) improve remission rates (O)?
2. AcquireSearch for best evidencePubMed, Cochrane Library, NICE guidelines
3. AppraiseCritically evaluate the evidenceCheck study design, bias, applicability
4. ApplyIntegrate evidence with clinical expertise and patient preferencesDiscuss options with patient considering their values
5. AssessEvaluate your practice and outcomesAudit remission rates in your patients

6.3 Levels of Evidence Pyramid

LevelStudy TypeStrength
1aSystematic review of RCTsStrongest
1bIndividual RCT with narrow CI
2aSystematic review of cohort studies
2bIndividual cohort study
3aSystematic review of case-control studies
3bIndividual case-control study
4Case series
5Expert opinionWeakest

6.4 NNT and NNH

MeasureFormulaInterpretation
NNT (Number Needed to Treat)1 / ARRNumber of patients needed to treat with intervention for one additional good outcome
NNH (Number Needed to Harm)1 / ARINumber of patients treated before one additional patient is harmed
ARR (Absolute Risk Reduction)CER − EERDifference in event rates between control and experimental groups

Example: If relapse rate is 40% with placebo and 25% with drug:

Clinical rule: Lower NNT = more effective treatment. Higher NNH = safer treatment. Compare NNT to NNH for risk-benefit.


7. META-ANALYSIS AND SYSTEMATIC REVIEW

7.1 Definitions

FeatureSystematic ReviewMeta-analysis
DefinitionStructured, reproducible synthesis of all available evidence on a questionStatistical pooling of results from multiple studies into a single summary estimate
StatisticsQualitative synthesis (narrative)Quantitative synthesis (statistical)
RelationshipEvery meta-analysis is within a systematic reviewNot every systematic review includes a meta-analysis
When meta-analysis is NOT doneStudies too heterogeneous, different outcomes, insufficient data

7.2 Steps in a Systematic Review / Meta-analysis

  1. Define research question (PICO)
  2. Protocol registration (PROSPERO)
  3. Comprehensive literature search (multiple databases: PubMed, Cochrane, Embase, PsycINFO)
  4. Study selection (inclusion/exclusion criteria, PRISMA flowchart)
  5. Data extraction (standardised forms)
  6. Quality assessment (Cochrane Risk of Bias tool for RCTs, Newcastle-Ottawa Scale for observational studies)
  7. Data synthesis (narrative or meta-analytic)
  8. Assess heterogeneity (I2, Q test)
  9. Assess publication bias (funnel plot, Egger's test)
  10. Report (PRISMA guidelines)

7.3 Forest Plot Interpretation

A forest plot displays:

How to read:

7.4 Heterogeneity

MeasureDefinitionInterpretation
Cochrane QChi-square test for heterogeneityp < 0.10 significant heterogeneity (liberal threshold because of low power)
I2Percentage of variability due to heterogeneity rather than chance0–25% = low, 25–50% = moderate, 50–75% = substantial, > 75% = considerable

If heterogeneity is high:

If heterogeneity is low:

7.5 Publication Bias

7.6 Network Meta-analysis (NMA)


8. RESEARCH METHODOLOGY

8.1 Research Question: PICO Framework

ComponentMeaningExample
PPopulationAdults with treatment-resistant depression
IInterventionKetamine infusion
CComparisonMidazolam (active placebo)
OOutcomeReduction in MADRS score at 24 hours

8.2 Hypothesis Types

TypeDefinitionExample
Null (H0)No difference or no associationThere is no difference in HAM-D scores between drug and placebo
Alternative (H1)There IS a difference or associationDrug group has lower HAM-D scores than placebo
Directional (one-tailed)Specifies the direction of differenceDrug group will have LOWER scores (not just different)
Non-directional (two-tailed)Difference exists but direction unspecifiedScores will DIFFER between groups

8.3 Sampling Methods

Probability Sampling (every member has a known chance of selection)
MethodHow It WorksAdvantageLimitation
Simple randomEvery individual has equal chance (lottery, random number table)Unbiased, representativeNeeds complete sampling frame
Stratified randomDivide population into strata (e.g., age groups), random sample from eachEnsures representation of subgroupsMore complex, needs knowledge of strata
ClusterDivide population into clusters (e.g., villages), randomly select entire clustersPractical for large, dispersed populationsHigher sampling error
SystematicEvery kth individual from a list (e.g., every 5th patient)Simple, quickBias if list has a periodic pattern
Non-Probability Sampling (not everyone has a chance of selection)
MethodHow It WorksUse Case
ConvenienceWhoever is availableQuick pilot studies
PurposiveResearcher selects based on characteristicsQualitative research, expert panels
SnowballParticipants recruit other participantsHard-to-reach populations (IV drug users, sex workers)
QuotaResearcher ensures certain proportions (like stratified but non-random)Market research

8.4 Sample Size Calculation

Why it matters: Too small Type II error (miss real effects). Too large wasteful, ethical concerns (unnecessary exposure).

Factors affecting sample size:

Factor · Effect on Sample Size
Effect size ↓ Smaller sample needed
Power ↑ (e.g., 0.80 0.90) ↑ Larger sample needed
Alpha level ↓ (e.g., 0.05 0.01) ↑ Larger sample needed
Variability ↑ Larger sample needed
Expected dropout ↑ Larger sample needed (inflate by 10–20%)

General formula concept: n = f(Zα, Zβ, σ, δ) where:

8.5 Bias Types

BiasDefinitionExampleHow to Minimise
Selection biasSystematic differences in who enters the studyOnly including patients from tertiary hospitalRandom sampling, clear inclusion criteria
Information biasSystematic errors in measuring exposure or outcomeRecall bias in case-control studiesStandardised instruments, blinding
Recall biasCases remember exposures differently than controlsMothers of children with birth defects recall medication use more accuratelyProspective design, objective records
Observer biasResearcher's expectations influence data collectionUnblinded rater scores drug group as improvedBlinding of assessors
Performance biasUnequal care between groupsIntervention group gets more attentionStandardised protocols, blinding
Attrition biasDifferential loss to follow-upSicker patients drop out of drug arm drug looks betterITT analysis, minimise dropout
Publication biasPositive results more likely publishedOnly significant antidepressant trials publishedTrial registries, grey literature search
ConfoundingThird variable associated with both exposure and outcomeCoffee drinkers have more lung cancer (confounder: smoking)Randomisation, restriction, matching, stratification, multivariate analysis
Lead-time biasEarlier detection inflates survival time without true benefitScreening detected cancer 2 years earlier but patient dies at same timeMeasure mortality, not survival from diagnosis
Berkson biasHospital-based studies create spurious associationsStudying depression and diabetes in hospitalised patients, both conditions increase hospitalisation probabilityCommunity-based studies

8.6 Confounding, Effect Modification, and Mediation

ConceptDefinitionHow to IdentifyHow to Handle
ConfoundingThird variable distorts the true associationCrude and adjusted estimates differRandomisation, stratification, multivariate analysis
Effect modificationEffect of exposure on outcome DIFFERS across levels of a third variableSignificant interaction termReport stratum-specific estimates (do NOT adjust away)
MediationThird variable lies on the causal pathway between exposure and outcomeExposure Mediator OutcomeMediation analysis (Baron & Kenny steps, Sobel test)

8.7 Odds Ratio: Detailed

From a 2×2 table:

Disease +Disease −
Exposedab
Unexposedcd

8.8 Qualitative Research Methods

MethodFocusApproachExample
PhenomenologyLived experience of a phenomenonIn-depth interviews, meaning analysisExperience of living with schizophrenia
Grounded theoryDevelop theory from dataIterative coding, constant comparison, theoretical samplingHow families cope with a member's addiction
EthnographyCulture and social behaviourImmersive observation in a communityMental health beliefs in a tribal community
Content analysisSystematic categorisation of textCoding themes from transcriptsThemes in suicide notes
Thematic analysisIdentifying patterns across data6-step Braun & Clarke methodBarriers to mental health treatment in rural India

Qualitative rigour criteria (equivalent to quantitative validity/reliability):

8.9 Factor Analysis


9. ADDITIONAL HIGH-YIELD CONCEPTS

9.1 Normal Distribution

9.2 Bonferroni Correction

9.3 Intention-to-Treat vs Per-Protocol

FeatureITTPer-Protocol
IncludesALL randomised participantsOnly those who completed protocol
PreservesRandomisation
BiasConservative (underestimates effect)May overestimate effect
Preferred forSuperiority trials (primary analysis)Complements ITT; preferred for non-inferiority trials

9.4 CONSORT, STROBE, PRISMA

GuidelineFor Which Study TypeKey Feature
CONSORTRCTsFlow diagram of participant progression
STROBEObservational studies (cohort, case-control, cross-sectional)22-item checklist
PRISMASystematic reviews and meta-analysesFlow diagram of study selection

KEY TABLES SUMMARY

Master Test Selection Table

ScenarioParametric TestNon-Parametric Alternative
2 independent groups, continuousIndependent t-testMann-Whitney U
2 paired groups, continuousPaired t-testWilcoxon signed-rank
3+ independent groups, continuousOne-way ANOVAKruskal-Wallis
3+ paired groups, continuousRepeated measures ANOVAFriedman test
2 categorical variablesChi-square test of independenceFisher's exact test
Correlation (continuous)Pearson rSpearman ρ
Predict continuous outcomeLinear regression
Predict binary outcomeLogistic regression

Cross-reference: D2 (Model Answers), D3 (Mnemonics), D4 (Comparisons), D6 (Quick Review)

Chapter 02

Model Answers



Q1. Define validity. Types. Importance in psychiatry. [10 marks: 2+6+2]

Exam Strategy

Split strictly: 1 para definition (2 marks), 4 types with examples (6 marks, 1.5 each), brief importance section (2 marks). Use a table for types.

Model Answer

Definition (2 marks)

Validity refers to the degree to which a measurement instrument accurately measures the construct it is intended to measure. A valid test measures what it claims to measure. For example, a valid depression scale should measure depression, not general distress or anxiety.

Types of Validity (6 marks)

TypeDefinitionExample
Face validityThe instrument appears to measure the intended construct on superficial examination. Assessed by non-experts.PHQ-9 "looks like" it measures depression, questions about mood, sleep, appetite. Weakest form of validity.
Content validityThe instrument covers the entire domain of the construct comprehensively. Assessed by expert panel.HAM-D includes items on depressed mood, guilt, insomnia, psychomotor changes, somatic symptoms, and suicidality, covering the full domain of depression.
Criterion validityThe instrument correlates with an established external standard (gold standard). Two subtypes: (a) Concurrent, new test and criterion measured simultaneously (e.g., comparing a new depression scale with HAM-D in same patients). (b) Predictive, test predicts a future outcome (e.g., MMSE score predicting progression to dementia at 2-year follow-up).
Construct validityThe instrument measures the theoretical construct. Two subtypes: (a) Convergent, correlates with measures of the same construct (BDI correlates with PHQ-9). (b) Discriminant, does not correlate with measures of different constructs (BDI does not correlate with mania rating scale).

Importance in Psychiatry (2 marks)


Q2. What is validity? Discuss types with examples. [10 marks: 3+7]: LONG ESSAY CANDIDATE

Exam Strategy

This is a long essay candidate, write 3+ pages. Expand each type with psychiatric examples. Include internal vs external validity as bonus types to demonstrate depth.

Model Answer

Definition of Validity (3 marks)

Validity is defined as the degree to which a test or measurement instrument accurately measures the construct it purports to measure. It addresses the fundamental question: "Is this test measuring what we think it is measuring?"

Validity is not a binary property but exists on a continuum. A test may be valid for one purpose but not another. For example, the MMSE is a valid screening tool for cognitive impairment but not a valid diagnostic tool for specific dementia subtypes.

The concept is central to psychiatry because, unlike other medical specialties, psychiatric diagnoses rely primarily on phenomenological assessment rather than laboratory investigations. The validity of our tools directly determines the quality of our clinical decisions.

Types of Validity (7 marks)

1. Face Validity

Face validity refers to whether an instrument appears, on the surface, to measure what it claims to measure. It is the weakest and most subjective form of validity, typically assessed by non-experts or the target population.

Example: The PHQ-9 has high face validity, questions about feeling down, sleep difficulties, poor appetite, and concentration problems are recognisable as depression symptoms to most patients and clinicians.

Limitation: Face validity can be misleading. A test may "look right" but actually measure a different construct. Conversely, some valid tests may lack face validity (e.g., projective tests like the Rorschach).

Relevance: Important for patient compliance, patients are more likely to complete a questionnaire that appears relevant to their problems.

2. Content Validity

Content validity refers to the degree to which the instrument comprehensively covers all aspects of the construct being measured. It is assessed by expert panels who evaluate whether items represent the full domain.

Example: The HAM-D (Hamilton Depression Rating Scale) has strong content validity because it includes items assessing depressed mood, feelings of guilt, suicidal ideation, insomnia (early, middle, late), work and activities, psychomotor retardation, psychomotor agitation, anxiety, somatic symptoms, loss of insight, and diurnal variation, covering the full syndromal domain of depression.

A scale that only measured mood and sleep would lack content validity for depression because it omits cognitive, psychomotor, and somatic dimensions.

Assessment method: Content Validity Index (CVI), proportion of experts rating each item as relevant.

3. Criterion Validity

Criterion validity is established when the instrument's scores correlate with an external criterion or gold standard. It has two subtypes:

(a) Concurrent validity: The test and criterion are measured at the same time. The new instrument is validated against an established measure.

Example: A newly developed Hindi depression scale is administered alongside the validated HAM-D to the same patients on the same day. If the correlation (r) is > 0.70, concurrent validity is established.

(b) Predictive validity: The test score predicts a future outcome.

Example: The AUDIT (Alcohol Use Disorders Identification Test) score at admission predicting development of alcohol withdrawal seizures during hospitalisation. A high AUDIT score predicting more severe withdrawal supports predictive validity.

Another example: GCS score predicting functional outcome at 6 months in traumatic brain injury.

4. Construct Validity

Construct validity examines whether the instrument truly measures the theoretical construct it claims to measure. This is the most rigorous and important form of validity. Two subtypes:

(a) Convergent validity: The instrument correlates positively with other measures of the same construct.

Example: The BDI-II shows high correlation (r = 0.85) with the PHQ-9, both measuring depression. This supports convergent validity of both instruments.

(b) Discriminant (divergent) validity: The instrument does NOT correlate with measures of theoretically unrelated constructs.

Example: The BDI-II should show low correlation with the YMRS (Young Mania Rating Scale). If it showed high correlation with a mania scale, its discriminant validity would be questionable, it might be measuring general distress rather than depression specifically.

Other methods: Known-groups validity (instrument distinguishes between groups known to differ, e.g., depression scale scores higher in clinically depressed patients than healthy controls), factorial validity (factor analysis confirms the hypothesised factor structure).

5. Internal Validity (bonus)

Internal validity refers to the degree to which the results of a study can be attributed to the intervention rather than confounding variables. It applies to study designs rather than instruments.

Example: A well-conducted double-blind RCT of an antidepressant has high internal validity because randomisation controls for confounders, and blinding prevents bias.

Threats: Selection bias, maturation, history, testing effects, attrition, regression to the mean.

6. External Validity (bonus)

External validity refers to the generalisability of study results to other populations, settings, and times.

Example: An RCT conducted exclusively in urban tertiary centres with highly selected patients may lack external validity for rural primary care settings.

Relevance to Indian psychiatry: Many psychiatric treatment guidelines are based on Western studies, their external validity for the Indian population (different pharmacogenomics, cultural factors, disease presentation) requires careful consideration.


Q3. Define reliability and validity. Significance. Types. [10 marks: 3+2+5]

Exam Strategy

Balanced coverage needed. Give equal weight to both concepts. Use a comparison table to demonstrate understanding.

Model Answer

Definitions (3 marks)

Reliability is the degree to which a measurement instrument produces consistent and reproducible results when applied repeatedly under similar conditions. A reliable instrument yields the same result each time it measures the same underlying construct, assuming the construct has not changed.

Validity is the degree to which a measurement instrument accurately measures the construct it is intended to measure. A valid instrument measures what it claims to measure.

The relationship between reliability and validity is asymmetric: reliability is a necessary but not sufficient condition for validity. A test can be reliable but not valid (consistently measuring the wrong thing), but a test cannot be valid without being reliable.

Significance (2 marks)

Types (5 marks)

Types of Reliability:

TypeWhat It MeasuresStatisticExample
Test-retestStability over timePearson r, ICCAdministering BPRS to same patients 2 weeks apart
Inter-raterAgreement between ratersCohen's kappaTwo psychiatrists rating same patient on PANSS
Split-halfInternal consistency (two halves)Spearman-Brown corrected rOdd-numbered items vs even-numbered items of BDI
Internal consistencyHow well items measure same constructCronbach's alphaAll 21 items of BDI measuring depression (α > 0.7)

Types of Validity:

TypeWhat It MeasuresExample
FaceSurface appearancePHQ-9 looks like a depression questionnaire
ContentDomain coverageHAM-D covers all depression dimensions
Criterion (concurrent + predictive)Correlation with gold standardNew scale correlates with HAM-D (concurrent); AUDIT predicts withdrawal severity (predictive)
Construct (convergent + discriminant)Theoretical construct accuracyBDI correlates with PHQ-9 (convergent); BDI does not correlate with YMRS (discriminant)

Q4. What is reliability? Types with examples. Internal and external validity. [10 marks: 2+4+4]

Exam Strategy

Heavy on reliability (6 marks total with definition). Separate section for internal and external validity with threats.

Model Answer

Definition (2 marks)

Reliability is the extent to which a measurement instrument yields consistent, reproducible, and dependable results across different occasions, raters, or forms. If the underlying construct has not changed, a reliable instrument will produce the same score on repeated measurement.

Types of Reliability with Examples (4 marks)

1. Test-retest reliability: Measures the stability of scores over time. The same instrument is administered to the same individuals at two different time points, and the correlation between scores is calculated.

Example: The BPRS is administered to 50 patients with schizophrenia at baseline and again 2 weeks later (assuming clinical stability). A Pearson correlation coefficient of r = 0.85 indicates good test-retest reliability.

Consideration: The time interval matters, too short risks practice effects; too long risks genuine change in the construct.

2. Inter-rater reliability: Measures agreement between two or more raters assessing the same subject. Quantified using Cohen's kappa (κ) for categorical data or Intraclass Correlation Coefficient (ICC) for continuous data.

Example: Two psychiatrists independently rate the same videotaped clinical interview using the PANSS. Cohen's kappa of 0.75 indicates substantial inter-rater agreement. This is critical for multisite clinical trials where different raters must produce comparable scores.

3. Split-half reliability: The test is divided into two halves (usually odd vs even items), and the correlation between halves is calculated. Corrected using the Spearman-Brown prophecy formula.

Example: The 21-item BDI is split into odd-numbered items and even-numbered items. A corrected correlation of r = 0.88 indicates good split-half reliability.

4. Internal consistency (Cronbach's alpha): Measures how well all items in a test measure the same underlying construct. Essentially an average of all possible split-half combinations.

Example: Cronbach's alpha of 0.89 for the PHQ-9 means the 9 items consistently measure the same construct (depression). An alpha > 0.70 is acceptable; > 0.90 may indicate item redundancy.

Internal and External Validity (4 marks)

Internal Validity

Internal validity is the degree to which a study's results accurately reflect the true relationship between the independent and dependent variables, free from confounding influences. It answers: "Did the intervention truly cause the observed effect?"

Threats to internal validity:

Example: A double-blind, randomised, placebo-controlled trial of a new antipsychotic has high internal validity, randomisation balances confounders, blinding prevents observer and performance bias, and placebo controls for non-specific effects.

External Validity

External validity is the degree to which study findings can be generalised to other populations, settings, and times.

Threats to external validity:

Example: The STAR*D trial enrolled patients from both primary care and psychiatric clinics across the US, enhancing external validity. However, it excluded patients with bipolar disorder and active substance dependence, limiting generalisability to those populations.

The tension: RCTs maximise internal validity (controlled conditions) but may sacrifice external validity (strict criteria create unrepresentative samples). Pragmatic trials attempt to balance both.


Q5. What is psychiatric epidemiology? Uses. Types of designs. [10 marks: 2+4+4]

Exam Strategy

Define psychiatric epidemiology, list uses with Indian examples, then cover 4 major study designs.

Model Answer

Definition (2 marks)

Psychiatric epidemiology is the study of the distribution, determinants, and outcomes of mental disorders in defined populations. It applies epidemiological methods to understand the frequency (incidence, prevalence), risk factors, natural history, and burden of psychiatric illnesses, and to evaluate the effectiveness of interventions.

It differs from clinical psychiatry in that its unit of analysis is the population rather than the individual patient.

Uses of Psychiatric Epidemiology (4 marks)

  1. Estimation of disease burden: Determines prevalence and incidence of mental disorders. The National Mental Health Survey (NMHS) 2016 estimated the lifetime prevalence of mental disorders in India at 13.7%, with a treatment gap of 70–92%.
  1. Identification of risk and protective factors: Identifying that childhood adversity increases risk of adult depression (ACE studies), or that social support is protective against PTSD after trauma.
  1. Planning mental health services: Prevalence data guides resource allocation. India's District Mental Health Programme (DMHP) expansion was informed by epidemiological data on service gaps.
  1. Evaluation of interventions: Community-based studies evaluating the effectiveness of DMHP, or the MANAS trial evaluating collaborative care for depression in primary care in Goa.
  1. Understanding natural history: Longitudinal studies tracking the course of first-episode psychosis inform treatment duration guidelines.
  1. Policy development: WHO's Mental Health Atlas, Global Burden of Disease data, and NMHS data inform the Mental Healthcare Act 2017 and National Mental Health Policy.

Types of Study Designs (4 marks)

1. Cross-sectional (Prevalence study)

2. Case-control study

3. Cohort study

4. Randomised Controlled Trial


Q6. Describe different types of epidemiological studies in Psychiatry. [10 marks]

Exam Strategy

Cover all major types. Use a comparison table. Add psychiatric examples for each.

Model Answer

Introduction

Epidemiological studies in psychiatry can be classified as observational (researcher observes without intervening) or experimental (researcher assigns an intervention). Each design has specific strengths, limitations, and appropriate applications.

Classification of Epidemiological Studies

A. Observational Studies

1. Descriptive Studies

(a) Case report/case series:

(b) Cross-sectional (Prevalence study):

(c) Ecological study:

2. Analytical Observational Studies

(a) Case-control study:

(b) Cohort study:

Comparison of Analytical Observational Designs:

FeatureCase-ControlCohort
DirectionBackwardForward
Starts withDisease statusExposure status
MeasureOdds RatioRelative Risk
Best forRare diseasesRare exposures
Bias riskRecall biasLoss to follow-up
CostLowerHigher

B. Experimental Studies

(a) Randomised Controlled Trial (RCT):

(b) Quasi-experimental study:

(c) Community trial:


Q7. Discuss Non-Parametric Tests. [10 marks]

Exam Strategy

Define, explain when to use, then describe individual tests with examples. Include comparison table.

Model Answer

Definition and Rationale

Non-parametric tests (also called distribution-free tests) are statistical tests that do not assume the data follows a specific probability distribution (typically the normal distribution). They are used when:

Non-parametric tests work on ranks rather than raw values, making them more robust to outliers and distribution violations.

Individual Non-Parametric Tests

1. Mann-Whitney U Test

2. Wilcoxon Signed-Rank Test

3. Kruskal-Wallis Test

4. Friedman Test

5. Spearman's Rank Correlation (ρ)

6. Fisher's Exact Test

7. Sign Test

Comparison Table: Parametric vs Non-Parametric Equivalents

ParametricNon-ParametricPurpose
Independent t-testMann-Whitney U2 independent groups
Paired t-testWilcoxon signed-rank2 related groups
One-way ANOVAKruskal-Wallis3+ independent groups
Repeated measures ANOVAFriedman3+ related groups
Pearson rSpearman ρCorrelation
Chi-squareFisher's exactCategorical data (small n)

Advantages of Non-Parametric Tests:

Limitations:


Q8. Chi Square test. [10 marks]

Exam Strategy

Define, explain types, give formula, worked example, assumptions, and psychiatric application.

Model Answer

Definition

The Chi-square (χ2) test is a non-parametric statistical test used to examine the association between two categorical (nominal or ordinal) variables. It compares observed frequencies with expected frequencies under the null hypothesis of no association.

Types of Chi-Square Test

1. Chi-Square Test of Independence

2. Chi-Square Goodness of Fit

Formula

χ2 = Σ [(O − E)2 / E]

Where: O = observed frequency, E = expected frequency

For a 2×2 table, E for each cell = (row total × column total) / grand total

Degrees of Freedom: df = (rows − 1) × (columns − 1). For a 2×2 table, df = 1.

Worked Example

A study examines whether gender is associated with treatment response:

RespondedDid not respondTotal
Male302050
Female401050
Total7030100

Expected frequencies (under H0):

χ2 = (30-35)2/35 + (20-15)2/15 + (40-35)2/35 + (10-15)2/15

χ2 = 0.71 + 1.67 + 0.71 + 1.67 = 4.76

With df = 1, critical value at α = 0.05 is 3.84. Since 4.76 > 3.84, p < 0.05. We reject H0, there is a significant association between gender and treatment response.

Assumptions and Conditions

  1. Data must be categorical (nominal or ordinal)
  2. Observations must be independent (each subject contributes to only one cell)
  3. Expected frequency ≥ 5 in at least 80% of cells
  4. No cell should have an expected frequency < 1
  5. If assumptions are violated use Fisher's exact test

Yates' Correction for Continuity

- Formula: χ2 = Σ [(O − E− 0.5)2 / E]

Psychiatric Applications


Q9. Paired t Test. [10 marks]

Exam Strategy

Define, explain when to use, assumptions, formula, worked example with psychiatric data, interpretation.

Model Answer

Definition

The paired t-test (also called dependent samples t-test) is a parametric statistical test used to compare the means of two related measurements from the same subjects or matched pairs. It tests whether the mean difference between paired observations is significantly different from zero.

When to Use

Assumptions

  1. Continuous data (interval or ratio scale)
  2. Normal distribution of differences (not the raw scores, the differences between pairs must be approximately normal)
  3. Random sampling from the population
  4. Paired observations, each subject has two measurements
  5. No significant outliers in the difference scores

Formula

t = d / (SD_d / √n)

Where:

Worked Example

A psychiatrist administers the HAM-D to 8 patients before and after 6 weeks of SSRI treatment:

PatientPre-treatmentPost-treatmentDifference (d)
124186
220146
328226
422202
5261610
618126
730246
824168

Calculations:

Since 7.85 > 2.365, p < 0.05. There is a statistically significant reduction in HAM-D scores after 6 weeks of SSRI treatment.

Interpretation

Advantages Over Independent t-test

When NOT to Use

Psychiatric Applications


Q10. Enumerate statistical tests. Describe t-test with example. [10 marks: 4+6]

Exam Strategy

First part: classify all tests in a table (4 marks). Second part: detailed t-test with both types and an example (6 marks).

Model Answer

Enumeration of Statistical Tests (4 marks)

CategoryParametricNon-Parametric
2 independent groupsIndependent t-testMann-Whitney U
2 paired groupsPaired t-testWilcoxon signed-rank
3+ independent groupsOne-way ANOVAKruskal-Wallis
3+ related groupsRepeated measures ANOVAFriedman test
CorrelationPearson rSpearman ρ
Categorical variablesChi-square, Fisher's exact
Prediction (continuous outcome)Linear regression
Prediction (binary outcome)Logistic regression

Additional tests: Two-way ANOVA (two factors), ANCOVA (adjusting for covariates), MANOVA (multiple dependent variables), McNemar's test (paired categorical), Cochran's Q (3+ paired categorical).

The t-test, Detailed Description (6 marks)

The t-test is a parametric statistical test that compares the means of one or two groups to determine if there is a statistically significant difference. Developed by William Sealy Gosset (published under the pseudonym "Student" in 1908).

General Assumptions:

Type 1: Independent (Unpaired) t-test

Purpose: Compares means of two independent groups.

Formula: t = (X1 − X2) / √(Sp2 × (1/n1 + 1/n2))

Where Sp2 is the pooled variance.

Example: Comparing mean HAM-D scores between 30 patients receiving fluoxetine and 30 patients receiving placebo after 8 weeks:

When to use Welch's t-test: If Levene's test is significant (unequal variances), use Welch's t-test which does not assume equal variances.

Type 2: Paired (Dependent) t-test

Purpose: Compares means of two related measurements (before-after, matched pairs).

Formula: t = d / (SD_d / √n)

Example: Measuring PANSS scores in 25 patients before and 4 weeks after starting risperidone:

Type 3: One-sample t-test

Purpose: Compares a sample mean to a known or hypothesised population mean.

Example: Testing whether the mean IQ of patients with treatment-resistant schizophrenia differs from the population mean of 100.

Reporting Format

"An independent samples t-test revealed a statistically significant difference in HAM-D scores between the fluoxetine group (M = 10.5, SD = 4.2) and the placebo group (M = 16.8, SD = 5.1), t(58) = −5.35, p < 0.001, Cohen's d = 1.35."


Q11. Enumerate statistical tests. Describe paired t-test. Two non-parametric tests. [10 marks: 4+2+4]

Exam Strategy

Quick enumeration table (4 marks). Concise paired t-test (2 marks). Two non-parametric tests in detail (4 marks).

Model Answer

Enumeration of Statistical Tests (4 marks)

See Q10 for complete table. Briefly:

Parametric tests: Independent t-test, paired t-test, one-way ANOVA, two-way ANOVA, repeated measures ANOVA, Pearson correlation, linear regression, logistic regression, ANCOVA.

Non-parametric tests: Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis, Friedman, Spearman correlation, Chi-square, Fisher's exact test, McNemar's test, Sign test, Cochran's Q.

Paired t-test (2 marks)

The paired t-test compares the means of two related measurements from the same subjects. It calculates the difference (d) for each pair, then tests whether the mean difference (d) is significantly different from zero.

Formula: t = d / (SD_d / √n), df = n − 1

Assumptions: Continuous data, normal distribution of differences, paired observations.

Example: Comparing BDI scores in 20 patients before and after 12 sessions of CBT. If mean difference = 8.5 (SD = 3.2), t = 8.5/(3.2/√20) = 11.87, p < 0.001, significant improvement after CBT.

Two Non-Parametric Tests (4 marks)

1. Mann-Whitney U Test (2 marks)

The Mann-Whitney U test is the non-parametric alternative to the independent samples t-test. It compares the distributions of two independent groups by ranking all observations from both groups together.

Procedure:

  1. Combine all observations and rank them from lowest to highest
  2. Sum the ranks for each group
  3. Calculate U statistic for each group
  4. Compare U to the critical value or calculate z-score for large samples

Assumptions: Ordinal or continuous data, independent observations, similar distribution shapes (tests whether one distribution is shifted relative to the other).

Example: Comparing satisfaction with care (5-point Likert scale: very dissatisfied to very satisfied) between patients in a community mental health centre (n = 25) and a tertiary psychiatry department (n = 30). Since Likert data is ordinal, the Mann-Whitney U test is appropriate rather than the t-test.

2. Wilcoxon Signed-Rank Test (2 marks)

The Wilcoxon signed-rank test is the non-parametric alternative to the paired t-test. It compares two related measurements by analysing the magnitude and direction of differences.

Procedure:

  1. Calculate the difference for each pair
  2. Rank the absolute differences (ignoring zeros)
  3. Assign the sign (+ or −) of the original difference to each rank
  4. Sum the positive ranks (W+) and negative ranks (W−)
  5. The test statistic is the smaller of W+ and W−

Assumptions: Ordinal or continuous data, paired observations, symmetric distribution of differences.

Example: Comparing self-rated anxiety on a 10-point visual analogue scale before and after a 20-minute relaxation exercise in 15 patients. Since the VAS data may not be normally distributed with a small sample, the Wilcoxon signed-rank test is more appropriate than a paired t-test.


Q12. Describe Factor Analysis and applications in psychiatric research. [10 marks]

Exam Strategy

Define, explain EFA vs CFA, key terms, procedure, then psychiatric applications (the bulk of the marks).

Model Answer

Definition

Factor analysis is a multivariate statistical technique that reduces a large number of observed variables into a smaller number of underlying latent factors (dimensions) that explain the patterns of correlations among the variables. It identifies clusters of intercorrelated variables that are driven by the same underlying construct.

Types

1. Exploratory Factor Analysis (EFA)

2. Confirmatory Factor Analysis (CFA)

Key Concepts

Term · Definition
Factor Latent (unobserved) variable underlying a group of observed variables
Factor loading Correlation between an observed variable and a factor (> 0.3 or > 0.4 considered meaningful)
Eigenvalue Amount of total variance explained by a factor
Kaiser criterion Retain factors with eigenvalue > 1
Scree plot Graph of eigenvalues; retain factors before the "elbow"
Communality Proportion of a variable's variance explained by extracted factors
Rotation Simplifies factor structure. Varimax = orthogonal (factors uncorrelated). Oblimin/promax = oblique (factors allowed to correlate)
KMO Kaiser-Meyer-Olkin measure of sampling adequacy (> 0.6 acceptable, > 0.8 good)
Bartlett's test Tests whether the correlation matrix is an identity matrix (should be significant)

Procedure of EFA

  1. Assess suitability: KMO > 0.6, Bartlett's test significant
  2. Extract factors: Principal Component Analysis (PCA) or Principal Axis Factoring
  3. Determine number of factors: Eigenvalue > 1 (Kaiser), scree plot, parallel analysis
  4. Rotate factors: Varimax (orthogonal) or oblimin (oblique)
  5. Interpret: Examine factor loadings, name factors based on loading pattern
  6. Report: Variance explained, factor loadings, communalities

Applications in Psychiatric Research

1. Scale Development and Validation

2. Symptom Dimension Research

3. Personality Research

4. Diagnostic Classification

5. Treatment Research


Q13. What is EBM? Steps. Practice with example. [10 marks: 2+2+6]: LONG ESSAY CANDIDATE

Exam Strategy

Long essay. After definition and steps, walk through a COMPLETE clinical example through all 5 steps. This is what earns the 6 marks.

Model Answer

Definition (2 marks)

Evidence-Based Medicine (EBM) is defined as "the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients" (Sackett et al., 1996). It integrates three pillars:

  1. Best available research evidence
  2. Clinical expertise of the practitioner
  3. Patient values and preferences

EBM does not mean blindly following evidence, it means integrating evidence with clinical judgment and the individual patient's context.

Five Steps (2 marks)

  1. ASK, Formulate a clear clinical question (PICO format)
  2. ACQUIRE, Search for the best available evidence
  3. APPRAISE, Critically evaluate the evidence for validity and applicability
  4. APPLY, Integrate evidence with clinical expertise and patient preferences
  5. ASSESS, Evaluate the outcome and the process

Practice with a Detailed Clinical Example (6 marks)

Clinical Scenario:

A 35-year-old man with moderate major depressive disorder (MDD) has had partial response to fluoxetine 20 mg for 8 weeks. You are considering augmentation strategies.

Step 1: ASK, Formulate the PICO question

Component · Detail
P (Population) Adults with MDD who have partially responded to an SSRI
I (Intervention) Augmentation with aripiprazole
C (Comparison) Augmentation with lithium
O (Outcome) Remission rate (HAM-D ≤ 7) at 8 weeks

PICO question: "In adults with MDD who have partially responded to an SSRI, does augmentation with aripiprazole compared to lithium lead to higher remission rates?"

Step 2: ACQUIRE, Search for evidence

Search strategy:

Key findings:

Step 3: APPRAISE, Critical appraisal

For the network meta-analysis:

Step 4: APPLY, Integrate evidence with clinical context

This illustrates that EBM does not dictate a single correct answer, the "best evidence" is integrated with clinical expertise and patient preferences.

Step 5: ASSESS, Evaluate outcomes


Q14. What is EBM? Practice in India? Limitations. [10 marks]

Exam Strategy

Quick definition, then focus on Indian context (unique section) and limitations.

Model Answer

Definition

Evidence-Based Medicine (EBM) is the integration of the best available research evidence with clinical expertise and patient values in making clinical decisions (Sackett, 1996). It follows 5 steps: Ask, Acquire, Appraise, Apply, Assess.

Practice of EBM in India

Current status:

Challenges in Indian context:

  1. Limited Indian evidence base: Most psychiatric RCTs are conducted in Western, educated, industrialised, rich, democratic (WEIRD) populations. Pharmacogenomic differences (e.g., CYP2D6 polymorphisms in South Asians), cultural factors, and disease presentation may differ.
  1. Access to evidence: Despite Sci-Hub usage, many clinicians lack institutional access to journals. Open-access initiatives (PubMed Central, IJPM) help but coverage is incomplete.
  1. Treatment gap: With a 70–92% treatment gap (NMHS 2016) and 0.3 psychiatrists per 100,000 population, implementing evidence-based protocols is constrained by workforce shortages.
  1. Cultural applicability: Psychotherapy protocols (CBT, DBT) developed in Western settings require cultural adaptation for Indian patients (e.g., family involvement, spiritual frameworks, idioms of distress).
  1. Resource constraints: Evidence-based treatments like clozapine require blood monitoring, not feasible in rural areas. Cost of newer medications limits implementation.
  1. Research quality: Many Indian psychiatric studies have methodological limitations, small sample sizes, single-centre designs, lack of randomisation.

Positive developments:

Limitations of EBM

  1. Evidence hierarchy bias: Over-reliance on RCTs devalues clinical experience, qualitative research, and patient narratives
  2. Publication bias: Published evidence is skewed toward positive results
  3. Lag time: Evidence takes years to translate into practice (17-year bench-to-bedside gap)
  4. Not all questions are answerable by RCTs: Ethical/practical constraints (e.g., randomising to childhood adversity)
  5. Individual variation: Evidence provides population-level estimates, the individual patient may respond differently
  6. Industry influence: Pharmaceutical industry funding can bias trial design, outcome selection, and reporting
  7. Cookbook medicine risk: EBM misapplied becomes rigid algorithm-following, ignoring clinical nuance
  8. Equity issues: Evidence generated from privileged populations may not serve marginalised communities

Q15. Discuss levels of evidence and role in EBM. [10 marks]

Exam Strategy

Present the full hierarchy, explain each level, then discuss their role in clinical decision-making.

Model Answer

Introduction

Levels of evidence refer to a hierarchical ranking system that grades the quality and reliability of evidence from clinical research. This hierarchy helps clinicians quickly assess the strength of evidence supporting a clinical decision.

Levels of Evidence Hierarchy

LevelStudy TypeDescription
1aSystematic review of RCTsPooled analysis of multiple RCTs with meta-analysis. Strongest evidence.
1bIndividual RCT with narrow CISingle well-designed RCT with adequate power and narrow confidence intervals.
1cAll-or-none studiesDramatic effects where all patients previously died but now survive (e.g., insulin for DKA).
2aSystematic review of cohort studiesPooled analysis of observational cohort studies.
2bIndividual cohort study or low-quality RCTSingle cohort study or an RCT with methodological limitations.
2cOutcomes researchEcological studies using large databases.
3aSystematic review of case-control studiesPooled analysis of case-control studies.
3bIndividual case-control studySingle well-designed case-control study.
4Case seriesReport of a series of patients without a control group.
5Expert opinionOpinion of respected authorities, based on clinical experience, descriptive studies, or reports of expert committees.

The Evidence Pyramid (Visual Description)

The evidence pyramid has its base as expert opinion (broadest, weakest) and its apex as systematic reviews/meta-analyses (narrowest, strongest). Moving up the pyramid:

Role of Levels of Evidence in EBM

1. Guiding treatment decisions

2. Developing clinical guidelines

3. Identifying evidence gaps

4. Clinical decision-making in practice

5. Limitations of the hierarchy


Q16. What is quantitative research? Features, strengths, limitations. [10 marks: 3+3+2+2]

Exam Strategy

Define with features (3+3 marks), then strengths and limitations (2+2 marks).

Model Answer

Definition (3 marks)

Quantitative research is a systematic empirical investigation that uses numerical data, statistical analysis, and mathematical models to test hypotheses, establish relationships between variables, and make generalisable inferences about populations. It follows the scientific method of hypothesis formulation, data collection, statistical testing, and interpretation.

The underlying philosophy is positivism, the belief that reality is objective, measurable, and can be understood through systematic observation and experimentation.

Features (3 marks)

  1. Numerical data: Variables are measured and expressed as numbers (e.g., HAM-D scores, age, number of hospitalisations)
  2. Hypothesis-driven: Begins with a clear, testable hypothesis (H0 and H1)
  3. Structured design: Follows predetermined protocols, study design, sampling, data collection, and analysis are planned before data collection begins
  4. Large sample sizes: Aims for adequate statistical power through sample size calculation
  5. Standardised instruments: Uses validated rating scales, structured interviews, laboratory measures
  6. Statistical analysis: Uses inferential statistics (t-test, ANOVA, regression) to test hypotheses
  7. Objectivity: Researcher maintains distance from participants; findings should be reproducible
  8. Generalisability: Aims to generalise findings from sample to population
  9. Control of variables: Attempts to control or adjust for confounders
  10. Replicability: Detailed methodology allows other researchers to replicate the study

Strengths (2 marks)

Limitations (2 marks)


Q17. Qualitative vs quantitative research. [10 marks: 2+5+3]

Exam Strategy

Define both (2 marks), detailed comparison table (5 marks), when to use each in psychiatry (3 marks).

Model Answer

Definitions (2 marks)

Quantitative research uses numerical data and statistical analysis to test hypotheses, measure variables, and generalise findings to populations. It follows a deductive approach (theory hypothesis observation).

Qualitative research uses non-numerical data (words, themes, narratives) to explore meaning, experience, and social processes. It follows an inductive approach (observation patterns theory).

Detailed Comparison (5 marks)

FeatureQuantitativeQualitative
PhilosophyPositivism (objective reality)Constructivism/interpretivism (subjective reality)
ApproachDeductive (theory-driven)Inductive (data-driven)
Data typeNumbers, measurementsWords, narratives, images
Sample sizeLarge (powered for statistics)Small (purposive, data saturation)
SamplingRandom/probabilityPurposive, snowball, theoretical
AnalysisStatistical (t-test, ANOVA, regression)Thematic, content, grounded theory, IPA
InstrumentsStandardised scales, questionnairesInterviews, focus groups, observation
Researcher roleDetached, objectiveEmbedded, reflexive
OutcomeGeneralisable findings, effect sizesRich descriptions, themes, theory
Question type"How much? How many? Is there a difference?""What is the experience like? How do people make sense of it?"
Rigour criteriaValidity, reliabilityCredibility, transferability, dependability, confirmability
CausationCan establish (RCTs)Cannot establish
ReplicationHigh (standardised methods)Low (context-dependent)

When to Use Each in Psychiatry (3 marks)

Quantitative is preferred when:

Qualitative is preferred when:

Mixed methods, integrating both:


Q18. What is meta-analysis vs systematic review? Key studies in psychiatry. [10 marks]

Exam Strategy

Define and compare (4 marks), process (3 marks), landmark psychiatric studies (3 marks).

Model Answer

Definitions and Comparison (4 marks)

FeatureSystematic ReviewMeta-analysis
DefinitionA structured, reproducible method of identifying, evaluating, and synthesising all available evidence relevant to a specific research questionA statistical technique that combines quantitative results from multiple studies into a single pooled estimate
NatureQualitative synthesisQuantitative synthesis
RelationshipMay or may not include a meta-analysisAlways part of a systematic review
When meta-analysis is NOT doneWhen studies are too heterogeneous, use different outcome measures, or have insufficient data for pooling
OutputNarrative synthesis with tablesForest plot, pooled effect size, I2, funnel plot

A meta-analysis is always nested within a systematic review, but not every systematic review performs a meta-analysis. When study heterogeneity is too high or studies are too clinically diverse, a narrative synthesis is more appropriate.

Process of Conducting a Systematic Review/Meta-analysis (3 marks)

  1. Formulate PICO question and register protocol (PROSPERO)
  2. Search strategy: Multiple databases (PubMed, Cochrane, Embase, PsycINFO), grey literature, hand-searching reference lists
  3. Study selection: Two independent reviewers screen titles/abstracts, then full texts; PRISMA flow diagram documents this
  4. Data extraction: Standardised forms extracting sample size, effect sizes, outcomes
  5. Quality assessment: Cochrane Risk of Bias tool (RCTs), Newcastle-Ottawa Scale (observational)
  6. Data synthesis:
  7. Fixed-effect model (assumes one true effect across studies)
  8. Random-effects model (assumes true effect varies, used when heterogeneity exists)
  9. Assess heterogeneity: I2 statistic, Cochrane Q test
  10. Assess publication bias: Funnel plot, Egger's test, trim-and-fill
  11. Sensitivity analysis: Removing one study at a time to test robustness
  12. Report: Following PRISMA guidelines

Forest plot interpretation: Each study represented by a square (point estimate) with horizontal line (CI). Diamond at bottom = pooled estimate. Vertical line of no effect (OR = 1 or MD = 0). If diamond does not cross this line overall result is statistically significant.

Key Meta-analyses and Systematic Reviews in Psychiatry (3 marks)

  1. Cipriani et al., Lancet 2018: Network meta-analysis of 21 antidepressants (522 trials, 116,477 patients). Found all antidepressants more effective than placebo. Amitriptyline, mirtazapine, and venlafaxine most effective; fluoxetine and escitalopram best accepted (lowest dropout).
  1. Leucht et al., Lancet 2012: Meta-analysis of antipsychotics for schizophrenia (65 RCTs). Found all antipsychotics more effective than placebo; effect sizes were moderate. Clozapine, amisulpride, olanzapine, and risperidone showed largest effects.
  1. Cuijpers et al. (multiple reviews): Extensive meta-analyses of psychotherapy for depression. CBT, IPT, and behavioural activation all effective; no consistent superiority of one modality over others (Dodo bird verdict). Combined psychotherapy + medication superior to either alone.
  1. Furukawa et al., World Psychiatry 2019: Network meta-analysis of initial treatment strategies for depression (pharmacotherapy vs psychotherapy vs combined). Combined treatment most effective.
  1. NICE guidelines for schizophrenia, depression, bipolar disorder: Based on systematic reviews of the evidence.
  1. Cochrane reviews in psychiatry: Cover topics from ECT for depression to clozapine for treatment-resistant schizophrenia to psychoeducation for bipolar disorder.

Q19. Odds ratio. Confounding, effect modification, mediation. [10 marks: 4+2+2+2]

Exam Strategy

Detailed OR section with 2×2 table (4 marks). Then concise but clear explanations of each concept (2 marks each).

Model Answer

Odds Ratio (4 marks)

Definition: The odds ratio (OR) is a measure of association between an exposure and an outcome in case-control and cross-sectional studies. It compares the odds of exposure in cases to the odds of exposure in controls.

2×2 Table:

Disease + (Cases)Disease − (Controls)
Exposedab
Unexposedcd

Formula: OR = (a × d) / (b × c)

Interpretation:

Example: A case-control study of childhood trauma and adult depression:

Depression (Cases)No Depression (Controls)
Trauma +8040
Trauma −2060

OR = (80 × 60) / (40 × 20) = 4800/800 = 6.0

Interpretation: Individuals with childhood trauma have 6 times the odds of developing depression compared to those without trauma.

OR vs RR: In case-control studies, we cannot calculate relative risk (because we select by outcome, not by exposure). OR approximates RR when the outcome is rare (< 10%).

Confounding (2 marks)

Definition: Confounding occurs when a third variable (confounder) is associated with both the exposure and the outcome, creating a spurious or distorted association between them.

Criteria for a confounder:

  1. Associated with the exposure
  2. Independently associated with the outcome
  3. Not on the causal pathway between exposure and outcome

Example: A study finds that coffee consumption is associated with lung cancer. However, coffee drinkers are more likely to smoke. Smoking is the confounder, it is associated with both coffee consumption and lung cancer. Once you adjust for smoking, the coffee-lung cancer association disappears.

How to control confounding:

Detection: If the adjusted (stratified or regression-adjusted) estimate differs from the crude estimate by > 10%, confounding is present.

Effect Modification (Interaction) (2 marks)

Definition: Effect modification occurs when the magnitude of the association between an exposure and an outcome differs across levels of a third variable (the effect modifier). Unlike confounding, effect modification is a real biological or social phenomenon and should be reported, not adjusted away.

Example: A study finds that SSRI treatment reduces depression scores by 8 points in women but only 3 points in men. Gender is an effect modifier, the treatment effect differs by gender.

How to detect:

How to handle: Report stratum-specific estimates. Do NOT adjust for effect modifiers (unlike confounders). If SSRIs work better in women, this is clinically important information.

Mediation (2 marks)

Definition: Mediation occurs when a third variable (mediator) lies on the causal pathway between the exposure and the outcome. The exposure causes the mediator, which in turn causes the outcome.

Causal pathway: Exposure Mediator Outcome

Example: Childhood trauma Maladaptive schemas Adult depression. Maladaptive schemas mediate the relationship between trauma and depression. Part of the effect of trauma on depression operates through the development of maladaptive schemas.

Testing mediation (Baron & Kenny, 1986):

  1. Exposure significantly predicts outcome (path c)
  2. Exposure significantly predicts mediator (path a)
  3. Mediator significantly predicts outcome when controlling for exposure (path b)
  4. The effect of exposure on outcome is reduced (partial mediation) or non-significant (full mediation) when mediator is in the model (path c')

Types:


Q20. Sampling techniques. Sample size. Principles of calculation. [10 marks: 2+2+6]

Exam Strategy

Quick sampling overview (2 marks), define sample size (2 marks), then detailed principles of calculation (6 marks).

Model Answer

Sampling Techniques (2 marks)

Probability sampling (every member has a known chance of selection):

Non-probability sampling (not everyone has equal chance):

Sample Size (2 marks)

Sample size is the number of participants required in a study to detect a clinically meaningful effect with adequate statistical power while maintaining a specified level of significance. An adequate sample size is the ethical minimum, too small wastes resources and cannot answer the question, too large exposes unnecessary participants to experimental conditions.

Principles of Sample Size Calculation (6 marks)

1. Type I Error Rate (α)

2. Statistical Power (1 − β)

3. Effect Size (δ)

4. Variability (σ)

5. Sample Size Formulas

For comparing two means (independent t-test):

n (per group) = 2 × [(Zα + Zβ)2 × σ2] / δ2

Where:

Worked example:

Comparing HAM-D scores between drug and placebo groups:

n = 2 × [(1.96 + 0.84)2 × 36] / 16 = 2 × [7.84 × 36] / 16 = 2 × 17.64 = 35.3

Need approximately 36 per group, 72 total.

For comparing two proportions:

n = [(Zα√(2pq) + Zβ√(p1q1 + p2q2))2] / (p1 − p2)2

6. Adjustment for Dropout

Adjusted n = n / (1 − expected dropout rate)

If expecting 20% dropout: 72 / 0.80 = 90 total

7. Other Considerations

8. Software and Resources


Cross-reference: D1 (Study Notes), D3 (Mnemonics), D4 (Comparisons), D6 (Quick Review)

Chapter 03

Mnemonics & Memory Tricks



Mnemonic 1: Types of Validity: "FaCe CoCo"

Fa = Face validity

Ce = Content validity (Experts judge domain coverage)

Co = Criterion validity (Concurrent + Predictive, against a gold standard)

Co = Construct validity (Convergent + Discriminant, theoretical construct)

Memory aid: "FaCe CoCo", like the face of a coconut shell. Face is surface level (weakest), and you crack through to get to the real content, criterion, and construct inside.

Expanded:


Mnemonic 2: Types of Reliability: "TISI"

T = Test-retest (Time stability)

I = Inter-rater (Individuals agree)

S = Split-half (Splitting the test)

I = Internal consistency (Items hang together, Cronbach's alpha)

Memory aid: "TISI" sounds like "Tissue", reliability is like a tissue, consistent and uniform throughout. If one part tears easily while the rest is strong, the tissue is not reliable.

Statistics to remember:


Mnemonic 3: Study Design Hierarchy: "Smart Researchers Can Create Excellent Cases"

From strongest to weakest:

Systematic review / Meta-analysis

Randomised Controlled Trial

Cohort study

Case-control study

Epidemiological cross-sectional study

Case series / Case report

(Expert opinion at the bottom)

Memory aid: "Smart Researchers Can Create Excellent Cases", and the smartest ones (systematic reviews) sit at the top.


Mnemonic 4: Parametric ↔ Non-Parametric Test Pairs: "I Must Pay Willingly, One Keeps Repeating Freely, Plus Spares"

ParametricNon-ParametricMnemonic Word
Independent t-testMann-Whitney U"I Must"
Paired t-testWilcoxon signed-rank"Pay Willingly"
One-way ANOVAKruskal-Wallis"One Keeps"
Repeated measures ANOVAFriedman"Repeating Freely"
Pearson rSpearman ρ"Plus Spares"

Memory aid: "I Must Pay Willingly, One Keeps Repeating Freely, Plus Spares", imagine paying your statistician willingly because they keep repeating analyses freely and always have spare tests.


Mnemonic 5: Five Steps of EBM: "5 A's"

Ask Formulate PICO question

Acquire Search for evidence

Appraise Critically evaluate evidence

Apply Integrate with expertise and patient values

Assess Evaluate outcomes

Memory aid: Already a built-in mnemonic. The 5 A's. Think "AAAAA", like the five-star rating you're giving to your evidence-based practice.


Mnemonic 6: Levels of Evidence: "Meta Rules Cohorts, Cases Come Last, Experts Guess"

Meta-analysis / Systematic review (Level 1a)

Randomised Controlled Trial (Level 1b)

Cohort study (Level 2)

Case-control study (Level 3)

Case series (Level 4)

Last: Expert opinion (Level 5)

Guess = expert opinion is basically an educated guess

Memory aid: "Meta Rules, Cohorts and Cases Come Last, Experts Guess", meta-analyses rule the hierarchy, and experts at the bottom are basically guessing (educated guessing, but still).


Mnemonic 7: Sensitivity/Specificity: "SnNOut / SpPIn"

SnNOut = Sensitive test, Negative result, rules Out disease

SpPIn = Specific test, Positive result, rules In disease

Memory aid: "SnNOut", if you're sick and the sensitive test says No, you're Out (no disease). "SpPIn", the Specific test says Positive, you're In (have the disease).

Why this works:

Bonus, the 2×2 table positions:


Mnemonic 8: Bias Types: "SCORES + PB"

Selection bias, who enters the study

Confounding, third variable distorts

Observer bias, researcher sees what they expect

Recall bias, cases remember differently

Equipment/Instrumentation bias, tools change

Survivor/Attrition bias, differential dropout

+PB:

Performance bias, unequal care between groups

Berkson bias, hospital-based studies create false associations

Memory aid: "SCORES + PB", think of "SCORES" like test scores that can be biased in many ways, plus "PB" for publication bias (which you should always mention in any meta-analysis answer).


Mnemonic 9: PICO Framework: "Patient, Intervention, Comparison, Outcome"

P = Patient/Population/Problem

I = Intervention (or Exposure)

C = Comparison (or Control)

O = Outcome

Memory aid: "PICO" already sounds like "pick-o", you PICK your research question elements. For observational studies, some use PECO (E = Exposure).

Quick template: "In [P], does [I] compared to [C] improve [O]?"


Mnemonic 10: Type I vs Type II Errors: "The Boy Who Cried Wolf"

Type I error (α) = False alarm = "Crying wolf when there is no wolf"

Type II error (β) = Missed wolf = "Not crying wolf when the wolf is real"

Memory aid: "Type I = I see it (but it's not there). Type II = Too blind to see it (but it IS there)."


Mnemonic 11: When to Use Which Test: "The 2-3 Rule"

2 groups?

3+ groups?

Categorical?

Correlation?

Decision flowchart in words:

  1. What's the outcome? Continuous or Categorical?
  2. If continuous: How many groups? 2 or 3+?
  3. If 2 groups: Paired or Independent?
  4. Normally distributed? Yes = parametric, No = non-parametric

Mnemonic 12: Threats to Internal Validity: "SHRIMP-T"

S = Selection bias

H = History (external events)

R = Regression to the mean

I = Instrumentation changes

M = Maturation (natural changes)

P = Practice/Testing effects

T = Testing attrition (dropout)

Memory aid: "A SHRIMP-T study has many internal validity threats", picture a tiny shrimp of a study, full of holes and threats. A well-designed RCT is the big fish that overcomes these.


Mnemonic 13: Cohen's Kappa Interpretation: "Slight Fair Moderate Substantial Perfect"

KappaAgreementMnemonic
< 0.20SlightSomeone barely trying
0.21–0.40FairFair attempt
0.41–0.60ModerateMiddle ground
0.61–0.80SubstantialSolid work
0.81–1.00Almost perfectAmazing

Memory aid: "Some Fair Midfielders Score Amazing goals", kappa goes from slight (barely trying) to almost perfect (amazing).


Mnemonic 14: Forest Plot Reading: "SLIDE"

S = Square = individual study point estimate (size = weight)

L = Line through square = confidence interval for that study

I = Invisible vertical line at null (OR=1 or MD=0) = line of no effect

D = Diamond at bottom = pooled/summary estimate (width = CI)

E = Effect is significant if diamond does NOT cross the line of no effect

Memory aid: "SLIDE", you slide your eyes down the forest plot from individual studies to the summary diamond.


Mnemonic 15: Qualitative Rigour: "CriTDeC"

Qualitative equivalents of quantitative concepts:

Cri = Credibility (≈ internal validity)

T = Transferability (≈ external validity)

De = Dependability (≈ reliability)

C = Confirmability (≈ objectivity)

Memory aid: "CriTDeC", say it fast, sounds like "critical deck", your critical deck of cards for evaluating qualitative research quality.


BONUS: Quick Number Anchors

What · Number to Remember
Acceptable Cronbach's alpha > 0.70
Substantial Cohen's kappa > 0.60
Standard p-value threshold 0.05
Standard power 0.80
Cohen's d: small/medium/large 0.2 / 0.5 / 0.8
Z for α = 0.05 (two-tailed) 1.96
Z for power = 0.80 0.84
AUC excellent range 0.80–0.90
Chi-square min expected count ≥ 5
Normal distribution: ±1 SD 68%
Normal distribution: ±2 SD 95%

Cross-reference: D1 (Study Notes), D2 (Model Answers), D4 (Comparisons), D6 (Quick Review)

Chapter 04

High-Yield Comparisons



Table 1: Parametric vs Non-Parametric Tests (with Paired Equivalents)

PurposeParametric TestNon-Parametric TestData Type
2 independent groupsIndependent t-testMann-Whitney UContinuous
2 paired/related groupsPaired t-testWilcoxon signed-rankContinuous
3+ independent groupsOne-way ANOVAKruskal-WallisContinuous
3+ related groupsRepeated measures ANOVAFriedman testContinuous
CorrelationPearson rSpearman ρContinuous
2×2 categoricalChi-square / Fisher's exactCategorical
Paired categorical (2×2)McNemar's testCategorical
3+ paired categoricalCochran's QCategorical

When to use parametric: Normal distribution, continuous data, adequate sample size (n > 30 per group), homogeneity of variance.

When to use non-parametric: Ordinal data, skewed distribution, small sample size, outliers, violated assumptions.

Key distinction: Parametric tests use raw values and assume a specific distribution; non-parametric tests use ranks and are distribution-free.


Table 2: Type I vs Type II Errors

FeatureType I Error (α)Type II Error (β)
DefinitionRejecting H0 when it is trueFailing to reject H0 when it is false
In plain languageFinding a difference when none existsMissing a real difference
AnalogyFalse alarm (crying wolf)Missed detection (wolf gets in)
ConsequenceAdopting an ineffective treatmentMissing an effective treatment
Controlled bySignificance level (α), conventionally 0.05Statistical power (1 − β), conventionally 0.80
Relationship to sample sizeFixed by α level↓ as sample size ↑ (more power)
Clinical impactPatient receives unnecessary/ineffective treatmentPatient misses a beneficial treatment
Multiple comparisonsRisk increases with more tests (Bonferroni correction)
Exam mnemonicType I = I made it up (false positive)Type II = Too blind to see it (false negative)
Which is worse?Context-dependent: worse in safety trials (approving a harmful drug)Context-dependent: worse in screening (missing a treatable cancer)

The trade-off: Reducing α (stricter threshold) decreases Type I error but increases Type II error (and vice versa). The only way to reduce BOTH is to increase sample size.


Table 3: Sensitivity vs Specificity vs PPV vs NPV

MeasureFormulaQuestion AnsweredAffected By
SensitivityTP / (TP + FN)Of all DISEASED, how many tested positive?Intrinsic test property
SpecificityTN / (FP + TN)Of all HEALTHY, how many tested negative?Intrinsic test property
PPVTP / (TP + FP)Of all TEST POSITIVES, how many truly have disease?Prevalence ↑ PPV ↑
NPVTN / (FN + TN)Of all TEST NEGATIVES, how many are truly healthy?Prevalence ↑ NPV ↓
Clinical RuleMeaningUse
SnNOutSensitive test + Negative result = rules OUT diseaseScreening (don't want to miss cases)
SpPInSpecific test + Positive result = rules IN diseaseConfirmation (don't want false positives)
MeasureHigh Value MeansClinical Application
LR+ (Sensitivity / (1-Specificity))Strong positive test performanceLR+ > 10 = very useful positive result
LR− ((1-Sensitivity) / Specificity)Strong negative test performanceLR− < 0.1 = very useful negative result

Critical insight: PPV and NPV depend on disease prevalence. Even a test with 99% sensitivity and 99% specificity will have poor PPV in a low-prevalence population because the absolute number of false positives will overwhelm true positives.


Table 4: Study Designs Comparison

FeatureCross-sectionalCase-ControlCohortRCT
DirectionSnapshot (no time direction)Retrospective (outcome exposure)Prospective or retrospective (exposure outcome)Prospective (intervention outcome)
What it measuresPrevalenceOdds Ratio (OR)Relative Risk (RR), IncidenceRelative Risk, NNT, NNH
Starting pointDefined populationCases + Controls selectedExposed + Unexposed groups identifiedParticipants randomised
Can establish temporalityNoNo (inferred)YesYes
Can establish causationNoNo (association only)Stronger associationYes (strongest)
Best forPrevalence estimation, generating hypothesesRare diseasesRare exposures, incidenceTesting interventions
Sample size neededModerateSmall–moderateLargeLarge
Time requiredShort (one point)Short–moderateLong (years of follow-up)Moderate–long
CostLowLow–moderateHighVery high
Key bias riskPrevalence bias, temporal ambiguityRecall bias, selection biasLoss to follow-up, confoundingHawthorne effect, strict inclusion low external validity
Psychiatry exampleNMHS 2016 prevalence surveyTrauma in BPD vs controlsDunedin cohort studySTAR*D, CATIE
Level of evidenceLower (descriptive)Level 3bLevel 2bLevel 1b
Ethical concernMinimalMinimalMinimal (observational)Cannot randomise harmful exposures

Table 5: Qualitative vs Quantitative Research

FeatureQuantitativeQualitative
PhilosophyPositivism (objective, measurable reality)Constructivism (subjective, multiple realities)
ApproachDeductive (theory hypothesis data)Inductive (data patterns theory)
Research question"How much? How many? Is there a difference?""What is the experience? How do people make sense of it?"
DataNumbers (scores, counts, measurements)Words (transcripts, field notes, narratives)
Sample sizeLarge (powered by calculation)Small (purposive, until data saturation)
SamplingProbability (random, stratified, cluster)Non-probability (purposive, snowball, theoretical)
Data collectionQuestionnaires, rating scales, structured interviewsIn-depth interviews, focus groups, observation
AnalysisStatistical (t-test, ANOVA, regression)Thematic analysis, grounded theory, IPA, content analysis
Researcher roleDetached, objective observerEmbedded, reflexive participant
Rigour criteriaValidity + ReliabilityCredibility + Transferability + Dependability + Confirmability
OutputEffect sizes, p-values, confidence intervalsThemes, narratives, theoretical frameworks
GeneralisabilityHigh (if representative sample)Low (context-dependent; transferability instead)
CausationCan establish (experimental designs)Cannot establish (explores meaning, not causation)
ReplicationHigh (standardised methods)Low (unique contexts)
StrengthsPrecision, generalisability, causal inferenceDepth, context, lived experience, "why"
LimitationsReductionism, misses meaning, decontextualisesNot generalisable, time-consuming, researcher bias
Psychiatry exampleRCT of SSRIs for depressionPhenomenological study of voice-hearing experience
Mixed methodsCan combine both: quantitative outcomes + qualitative process evaluation (e.g., MANAS trial)

Table 6: Systematic Review vs Meta-analysis vs Network Meta-analysis

FeatureSystematic ReviewMeta-analysisNetwork Meta-analysis
DefinitionStructured, reproducible synthesis of all evidence on a questionStatistical pooling of results from multiple studies into a single estimateExtension of meta-analysis comparing multiple treatments simultaneously
Nature of synthesisQualitative (narrative)Quantitative (statistical)Quantitative (statistical, complex)
Includes statistical poolingNot necessarilyYes (always)Yes (always)
Types of evidence comparedAll studies on one questionDirect comparisons (A vs B)Direct AND indirect comparisons (A vs B, B vs C infer A vs C)
Key outputSummary tables, quality assessment, narrative conclusionPooled effect size, forest plot, I2Treatment rankings, league table, network diagram
Heterogeneity assessmentDescribed qualitativelyI2, Cochrane QI2, inconsistency, node-splitting
Publication biasDiscussed qualitativelyFunnel plot, Egger's testComparison-adjusted funnel plot
When NOT to poolWhen heterogeneity is extremeWhen transitivity assumption is violated
Key assumptionComprehensive search, reproducible methodsStudies are sufficiently similar to poolTransitivity (similar populations across comparisons)
Reporting guidelinePRISMAPRISMAPRISMA-NMA
Psychiatry exampleCochrane review of CBT for depressionCipriani et al., 2018 (antidepressants)Cipriani et al., 2018 (21 antidepressants ranked)
Level of evidence1a1a1a

Table 7: Probability vs Non-Probability Sampling

FeatureProbability SamplingNon-Probability Sampling
DefinitionEvery member of the population has a known, non-zero chance of selectionSelection is based on availability, judgment, or other non-random criteria
BasisRandom selectionConvenience, judgment, or self-selection
GeneralisabilityHigh (representative of population)Low (may not represent population)
Bias riskLower (randomisation minimises selection bias)Higher (selection bias inherent)
Sampling frameRequired (complete list of population)Not required
Cost and effortHigher (need complete frame, randomisation)Lower (practical, quick)
Use in researchQuantitative studies, surveys, clinical trialsQualitative studies, pilot studies, hard-to-reach populations
TypeMethodExample
Simple randomEvery individual has equal chanceRandom number table to select patients from hospital register
StratifiedDivide into strata, random sample from eachStratify by age group, randomly sample within each stratum
ClusterRandomly select entire clustersRandomly select 10 PHCs from a district, study all patients at those PHCs
SystematicEvery kth individualEvery 5th patient from the OPD register
ConvenienceWhoever is availablePatients attending OPD on the day of data collection
PurposiveResearcher selects based on characteristicsSelecting only patients with treatment-resistant depression for qualitative interviews
SnowballParticipants recruit othersStudying IV drug users, each participant refers peers
QuotaNon-random but ensures proportionsEnsuring 50% male and 50% female in sample (non-randomly selected within each group)

Table 8: Types of Bias in Research

Bias TypeDefinitionStageExamplePrevention
Selection biasSystematic difference between study participants and target populationDesignOnly enrolling tertiary hospital patients (sicker, more complex)Random sampling, clear inclusion criteria
Recall biasCases recall exposure differently from controlsData collectionMothers of children with defects remember medication use more accuratelyProspective design, objective records
Observer/Detection biasAssessor's expectations influence measurementData collectionUnblinded rater scores drug group as improvedBlinding of assessors
Performance biasDifferential treatment between groups (beyond the intervention)ConductIntervention group receives more clinician attentionDouble blinding, standardised protocols
Attrition biasDifferential loss to follow-up between groupsFollow-upSicker patients drop out of drug armITT analysis, minimise dropout
Publication biasStudies with positive results are more likely publishedReportingOnly significant antidepressant trials in literatureTrial registries, grey lit search, funnel plot
ConfoundingThird variable associated with both exposure and outcomeAnalysisCoffee lung cancer (confounder: smoking)Randomisation, stratification, multivariate analysis
Lead-time biasEarlier detection inflates apparent survival timeInterpretationScreening detects cancer 2 years earlier without changing mortalityMeasure mortality, not survival from diagnosis
Berkson biasSpurious association in hospital-based studiesDesignDepression-diabetes association inflated because both increase hospitalisationPopulation-based studies
Information biasSystematic errors in data collectionData collectionUsing an unvalidated questionnaireStandardised, validated instruments
Reporting biasSelective reporting of outcomesReportingReporting only the outcome that was significantPre-registration, protocol publication
Hawthorne effectParticipants change behaviour because they are observedConductPatients adhere better during a trial than in routine careNaturalistic designs, long follow-up
Volunteer biasVolunteers differ systematically from non-volunteersDesignVolunteers are healthier, more motivated, more educatedRandom sampling (when possible)

Cross-reference: D1 (Study Notes), D2 (Model Answers), D3 (Mnemonics), D6 (Quick Review)

Chapter 05

PYQ Frequency Analysis


Exam Pearl

Source: PG exams Dec 2011, Jun 2025 + PG exams 2013-2022


Executive Summary

Statistics and research methodology is a consistently tested cluster, arguably the most predictable in Paper I after neurotransmitters and sleep. The questions follow a very stable pattern: validity/reliability, study designs, statistical tests, and evidence-based medicine rotate reliably.

Key insight: Unlike clinical topics, statistics questions have almost no variability. The same 5-6 question templates recycle. Master those templates and you're covered.


Topic-Level Frequency

TopicExam MentionsAvg per ExamVerdict
Validity & Reliability5~0.18Every 5-6 exams
Epidemiological study designs4~0.14Every 6-7 exams
Statistical tests (t-test, Chi-square, ANOVA, non-parametric)5~0.18Every 5-6 exams
Evidence-based medicine5~0.18Every 5-6 exams
Qualitative vs Quantitative research3~0.11Every 8-9 exams
Meta-analysis / Systematic review3~0.11Emerging topic
Sampling & Sample size2~0.07Occasional
Combined cluster~27~0.96~1 question per exam

Key PYQs Identified

Validity & Reliability

  1. "Define validity. Name different types. Importance in psychiatry." [10 marks, split 2+6+2]
  2. "What is validity? Enumerate types. Criterion validity." [10 marks, split 2+3+5]
  3. "What is validity? Discuss different types with examples." [10 marks, split 3+7]
  4. "Define reliability and validity. Significance. Types." [10 marks, split 3+2+5]
  5. "What is reliability? Types with examples. Internal and external validity." [10 marks, split 2+4+4]

Study Designs & Epidemiology

  1. "What is psychiatric epidemiology? Uses. Types of designs." [10 marks, split 2+4+4]
  2. "Describe different types of epidemiological studies in Psychiatry." [10 marks]
  3. "What is genetic epidemiology? Three types of genetic studies." [10 marks]
  4. "Define epidemiology. Different types of epidemiological studies." [10 marks]

Statistical Tests

  1. "Discuss Non-Parametric Tests." [10 marks]
  2. "Chi Square test." [10 marks]
  3. "Paired t Test." [10 marks]
  4. "Enumerate statistical tests to measure differences. Describe t-test with example." [10 marks, split 4+6]
  5. "Describe Factor Analysis and its applications in psychiatric research." [10 marks]
  6. "Enumerate statistical tests. Describe paired t-test. Two non-parametric tests for group difference." [10 marks, split 4+2+4]

Evidence-Based Medicine

  1. "What is evidence based medicine? How is it practiced? Steps with example." [10 marks, split 2+2+6]
  2. "What is EBM? How is it practiced in India? Limitations." [10 marks]
  3. "Discuss levels of evidence and their role in EBM." [10 marks]
  4. "What is EBM? What is personalized medicine? Pharmacogenetics." [10 marks, split 2+3+5]

Research Methodology

  1. "What is quantitative research? Essential features, strengths, limitations." [10 marks]
  2. "What is qualitative research? How does it differ from quantitative?" [10 marks, split 2+5+3]
  3. "What is meta-analysis? Systematic review? Key meta-analytical studies in psychiatry." [10 marks]
  4. "Odds ratio. Confounding, effect modification, mediation." [10 marks, split 4+2+2+2]
  5. "How to generate research question? Types of hypothesis. Meta-analysis vs network analysis." [10 marks]
  6. "Sampling techniques. Sample size. Principles of sample size calculation." [10 marks, split 2+2+6]

Long Essay Candidates

RankTopicProbability
1"Define validity and reliability. Types. Clinical importance."Very High
2"What is EBM? Steps. Practice in India. Limitations."Very High
3"Enumerate statistical tests. Describe t-test/chi-square with examples."High
4"Describe different epidemiological study designs in psychiatry."High
5"Qualitative vs quantitative research, features, strengths, limitations."Medium-High

Exam Strategy

Must-Prepare (will be asked)

  1. Validity, define, 4 types (face, content, criterion [concurrent + predictive], construct [convergent + discriminant]), examples from psychiatry
  2. Reliability, define, types (test-retest, inter-rater, split-half, internal consistency/Cronbach's alpha), relationship to validity
  3. Statistical tests, when to use each: t-test (2 means, parametric), ANOVA (>2 groups), Chi-square (categorical data), Mann-Whitney/Wilcoxon (non-parametric equivalents)
  4. EBM, define, 5 steps (ask, acquire, appraise, apply, assess), levels of evidence pyramid

Should-Prepare

  1. Study designs, cross-sectional, case-control, cohort, RCT, ecological. Strengths/limitations table
  2. Meta-analysis, define, steps, forest plot, heterogeneity (I2), advantages over narrative review
  3. Sensitivity/Specificity/PPV/NPV, 2x2 table, formulae, clinical application

Nice-to-Know

  1. Factor analysis (exploratory vs confirmatory)
  2. Odds ratio vs relative risk
  3. Confounding and effect modification
  4. Sampling techniques (probability vs non-probability)


Analysis based on PG exams Dec 2011, Jun 2025 + PG exams 2013-2022.

Chapter 06

Quick Review



RECALL Questions (1–12)

Q1. Define face validity and give one limitation.

A: Face validity is the degree to which a test appears, on the surface, to measure what it intends to measure. It is assessed subjectively by non-experts. Limitation: It is the weakest form of validity, a test can look right but actually measure a different construct (e.g., a test that appears to measure anxiety might actually be measuring general distress).

Q2. What are the two subtypes of criterion validity?

A: (1) Concurrent validity, the test and the gold standard criterion are measured at the same time (e.g., a new depression scale administered alongside the HAM-D). (2) Predictive validity, the test predicts a future outcome (e.g., AUDIT score at admission predicting withdrawal severity during hospitalisation).

Q3. Name the four types of reliability and their associated statistics.

A:

  1. Test-retest Pearson r or ICC
  2. Inter-rater Cohen's kappa (κ)
  3. Split-half Spearman-Brown corrected r
  4. Internal consistency Cronbach's alpha (α)

Q4. What is the non-parametric alternative to the independent t-test?

A: The Mann-Whitney U test. It compares the rank distributions of two independent groups and does not assume normal distribution.

Q5. What is the non-parametric alternative to the paired t-test?

A: The Wilcoxon signed-rank test. It compares two related measurements by analysing the magnitude and direction of differences between pairs.

Q6. State the formula for sensitivity and specificity.

A: From a 2×2 table (a = TP, b = FP, c = FN, d = TN):

Q7. What are the 5 steps of EBM?

A: Ask (formulate PICO question), Acquire (search for evidence), Appraise (critically evaluate), Apply (integrate with expertise and patient values), Assess (evaluate outcomes). The "5 A's."

Q8. What is the difference between a systematic review and a meta-analysis?

A: A systematic review is a structured, reproducible synthesis of all available evidence on a research question (qualitative/narrative). A meta-analysis is the statistical pooling of results from multiple studies into a single quantitative estimate. Every meta-analysis is part of a systematic review, but not every systematic review includes a meta-analysis (e.g., when studies are too heterogeneous to pool).

Q9. Define NNT and give its formula.

A: NNT (Number Needed to Treat) = the number of patients who need to be treated with the intervention for one additional patient to benefit compared to the control. Formula: NNT = 1 / ARR, where ARR (Absolute Risk Reduction) = Control Event Rate − Experimental Event Rate. Lower NNT = more effective treatment.

Q10. What is the PICO framework?

A: A structured format for clinical research questions: P = Population/Patient, I = Intervention, C = Comparison/Control, O = Outcome. Example: "In adults with MDD (P), does CBT + SSRI (I) compared to SSRI alone (C) improve remission rates (O)?"

Q11. Name three probability and three non-probability sampling methods.

A: Probability: Simple random, stratified random, cluster sampling. Non-probability: Convenience, purposive, snowball sampling.

Q12. What does an I2 of 75% mean in a meta-analysis?

A: I2 = 75% indicates substantial heterogeneity, 75% of the variability across study results is due to real differences between studies (heterogeneity) rather than chance. This suggests a random-effects model should be used, and sources of heterogeneity should be explored through subgroup analysis or meta-regression. (Interpretation: 0–25% low, 25–50% moderate, 50–75% substantial, >75% considerable.)


APPLICATION Questions (13–25)

Q13. A study finds p = 0.03 with a 95% CI of 1.2–3.4 for the odds ratio. Interpret.

A: The result is statistically significant (p < 0.05). The odds of the outcome are 1.2 to 3.4 times higher in the exposed group compared to the unexposed group (95% confidence). Since the CI does not include 1 (the null value for OR), this confirms statistical significance. The point estimate of the OR lies at some value between 1.2 and 3.4 (likely around 2.0). The exposure is a significant risk factor for the outcome.

Q14. A researcher wants to compare anxiety scores (normally distributed) before and after a 12-week yoga intervention in the same 40 patients. Which test?

A: Paired t-test. Rationale: The data is continuous, normally distributed, and involves two measurements from the same subjects (before-after design). If normality of differences were violated, use Wilcoxon signed-rank test instead.

Q15. A study compares treatment response (responder/non-responder) across three drug groups. Which test?

A: Chi-square test of independence. Rationale: The outcome is categorical (responder vs non-responder) and there are three independent groups. This creates a 3×2 contingency table. If any expected cell count is < 5, use Fisher's exact test or collapse categories.

Q16. You want to compare depression scores (skewed distribution) across four therapy groups with 12 patients each. Which test?

A: Kruskal-Wallis test. Rationale: Four independent groups, continuous but non-normally distributed data, relatively small sample sizes. This is the non-parametric alternative to one-way ANOVA. If significant, follow up with Dunn's test for pairwise comparisons.

Q17. A case-control study shows OR = 4.5 (crude) for the association between childhood trauma and BPD. After adjusting for parental substance use, the OR becomes 2.1. What happened?

A: Confounding was present. Parental substance use was a confounder, it was associated with both the exposure (childhood trauma) and the outcome (BPD). The crude OR of 4.5 was inflated by the confounding effect. The adjusted OR of 2.1 represents the true association after removing the confounding effect of parental substance use. Since the crude and adjusted estimates differ substantially (> 10% change), confounding is confirmed.

Q18. A screening test has sensitivity = 95% and specificity = 90%. In a population with 1% prevalence of the disease, calculate the approximate PPV.

A: For 10,000 people: 100 have disease, 9,900 do not.

Despite excellent sensitivity and specificity, the PPV is only 8.8% because the disease is rare. Most positive results are false positives. This is why screening for low-prevalence conditions generates many false alarms.

Q19. In an RCT, 60% of the control group relapsed and 40% of the drug group relapsed. Calculate the NNT.

A:

Interpretation: You need to treat 5 patients with the drug to prevent one additional relapse compared to control. This is a clinically meaningful NNT.

Q20. A forest plot shows 8 studies. 5 studies have confidence intervals that cross the line of no effect. The summary diamond does NOT cross the line of no effect. Is the overall result significant?

A: Yes, the overall result is statistically significant. While individual studies may be non-significant (their CIs cross the null line), the pooled estimate (summary diamond) does not cross the line of no effect, indicating that when all evidence is combined, the treatment effect is statistically significant. This demonstrates the power of meta-analysis, pooling data from multiple underpowered individual studies can reveal a significant effect.

Q21. A researcher measures the correlation between BDI scores and hours of sleep per night in 200 patients with depression. BDI scores are normally distributed. Which test?

A: Pearson's correlation coefficient (r). Rationale: Both variables are continuous (BDI score and hours of sleep), data is normally distributed, and the researcher is examining a linear association between two variables. If either variable were non-normally distributed or ordinal, use Spearman's ρ instead.

Q22. A study of an SSRI vs placebo reports: "Treatment group improved by 3 points more on HAM-D (p = 0.04, Cohen's d = 0.15)." Is this clinically meaningful?

A: Likely not clinically meaningful despite statistical significance. While p = 0.04 is statistically significant, Cohen's d = 0.15 indicates a very small effect size (small = 0.2, medium = 0.5, large = 0.8). The 3-point HAM-D difference is at the borderline of the minimum clinically important difference (MCID, typically 3–4 points). This illustrates that statistical significance does not equal clinical significance, with a large enough sample, even trivially small differences become statistically significant. Clinical decision-making should weigh effect size and MCID, not just p-values.

Q23. A researcher wants to study the experience of living with treatment-resistant depression from the patient's perspective. Which research methodology is most appropriate?

A: Qualitative phenomenological study using in-depth semi-structured interviews. Phenomenology focuses on the lived experience of a phenomenon and aims to understand its essential meaning from the participant's perspective. Data would be analysed using interpretative phenomenological analysis (IPA) or Colaizzi's method. The sample would be small (8–15 participants), purposively selected, and data collection would continue until thematic saturation.

Q24. An RCT randomises 200 patients to drug vs placebo. During the trial, 30 patients in the drug arm stop taking the medication due to side effects. How should they be analysed?

A: Under Intention-to-Treat (ITT) analysis, all 200 patients are analysed in their originally assigned groups, the 30 who discontinued are still counted in the drug arm. This preserves the benefits of randomisation and provides a conservative estimate of treatment effect. ITT reflects real-world effectiveness (not all patients comply). A per-protocol analysis (excluding non-compliers) can be reported as secondary analysis but may overestimate the true treatment effect by excluding those who had side effects.

Q25. A funnel plot in a meta-analysis shows asymmetry, with missing studies in the lower-left region. What does this suggest?

A: This suggests publication bias, smaller studies with negative or null results (lower-left region = small sample size + small/negative effect) are missing from the literature, likely because they were never published. This means the pooled estimate may overestimate the true treatment effect. The researcher should: (1) run Egger's test to confirm statistically, (2) apply trim-and-fill method to adjust the estimate, (3) search for grey literature and unpublished data.


ANALYSIS Questions (26–35)

Q26. Why might a highly sensitive test have poor PPV in a low-prevalence population?

A: Because PPV depends on both the test's specificity and the disease prevalence. In a low-prevalence population, the vast majority of people are disease-free. Even with high sensitivity (catching nearly all true cases), the small number of true positives is overwhelmed by false positives from the large healthy population. For example, if sensitivity = 99% and specificity = 95%, and prevalence = 0.1%: in 100,000 people, 100 have disease (99 detected) but 4,995 healthy people test falsely positive. PPV = 99/(99+4,995) = 1.9%. The absolute number of false positives is determined by (1-specificity) × (number of healthy people), which is enormous when prevalence is low.

Q27. A researcher reports that Cohen's kappa for their diagnostic interview is 0.45. Is this adequate for clinical use? Why or why not?

A: κ = 0.45 indicates moderate agreement, which is generally not adequate for clinical use, especially for high-stakes decisions. In clinical diagnostics, substantial to almost perfect agreement (κ > 0.60, ideally > 0.80) is expected. A κ of 0.45 means raters disagree on a substantial proportion of cases even after accounting for chance agreement. This could lead to misdiagnosis depending on which clinician the patient sees. The instrument or training protocol needs improvement before clinical deployment. However, for screening purposes (lower stakes, further assessment follows), moderate reliability may be acceptable.

Q28. An RCT of a new antidepressant excludes patients with comorbid substance use, personality disorders, and suicidal ideation. How does this affect the study's validity?

A: This creates a tension between internal and external validity:

Q29. A researcher finds a statistically significant correlation (r = 0.15, p = 0.01) between social media use and depression scores in a sample of 5,000 adolescents. What should you conclude?

A: The finding is statistically significant but clinically trivial. r = 0.15 is a very weak correlation, it explains only 2.25% of the variance in depression scores (r2 = 0.0225). The statistical significance is driven entirely by the large sample size (n = 5,000), which gives power to detect even negligible effects. The 97.75% of variance is explained by other factors. Additionally, this is a correlation, not causation, it could be that (a) social media causes depression, (b) depression leads to more social media use, or (c) a third variable (e.g., loneliness, sleep deprivation) drives both. This highlights why effect size and clinical significance matter more than p-values.

Q30. Why is ITT analysis considered more conservative than per-protocol analysis in a superiority trial?

A: ITT analysis includes ALL randomised participants in their assigned groups, regardless of whether they completed the intervention or adhered to the protocol. This is conservative because:

  1. Non-compliers dilute the treatment effect, patients who stopped the drug (due to side effects, lack of response, etc.) are still counted in the drug group, pulling the drug group's outcome toward the control group's outcome.
  2. Preserves randomisation, maintaining the original randomised groups ensures that known and unknown confounders remain balanced.
  3. Reflects real-world effectiveness, in practice, not all patients comply, so ITT gives a more realistic estimate.

Per-protocol analysis, by excluding non-compliers, creates a biased sample (those who tolerated and adhered may be inherently different) and can overestimate the treatment effect. However, in non-inferiority trials, per-protocol is actually the primary analysis because ITT (by diluting differences) can falsely support non-inferiority.

Q31. A new psychiatric rating scale has a Cronbach's alpha of 0.97. Is this a problem?

A: Potentially, yes. While α > 0.70 indicates acceptable internal consistency, an α of 0.97 suggests item redundancy, the items are so highly intercorrelated that they are essentially asking the same thing in slightly different ways. This means the scale is longer than necessary without adding new information. It increases respondent burden, completion time, and the risk of response fatigue without improving measurement. The scale developer should consider reducing items through item-total correlation analysis, removing items with very high inter-item correlations (> 0.90) that are redundant. An optimal Cronbach's alpha is typically 0.80–0.90.

Q32. A case-control study finds that patients with schizophrenia are more likely to have been born in winter months. What type of bias could explain this finding, and what design would strengthen the evidence?

A: This could be a genuine finding (season of birth effect is well-replicated) but could also be affected by Berkson bias (if controls were hospital-based with different seasonal admission patterns) or selection bias (if the sampling frame introduced seasonal artefacts). Recall bias is less relevant here since date of birth is an objective fact from records. To strengthen the evidence: (1) A large prospective birth cohort study following individuals from birth would establish temporality without selection bias. (2) Using population-based controls rather than hospital controls eliminates Berkson bias. (3) The ecological fallacy should be considered if aggregate population-level birth data is used. The Scandinavian birth register studies (e.g., Danish cohort) provide the strongest evidence for this association.

Q33. Why is the random-effects model preferred over the fixed-effect model when heterogeneity is high in a meta-analysis?

A: The fixed-effect model assumes there is ONE true effect size and all variation between studies is due to sampling error (chance). The random-effects model assumes the true effect varies between studies (because of differences in populations, interventions, settings) and accounts for both within-study and between-study variance.

When heterogeneity is high (I2 > 50%), the fixed-effect assumption is violated, studies clearly differ in their true effects. Using a fixed-effect model would produce inappropriately narrow confidence intervals (false precision) because it ignores the real variability between studies. The random-effects model produces wider, more honest confidence intervals that reflect the genuine uncertainty. It also weights studies more equally (small studies get relatively more weight compared to fixed-effect), which can be both an advantage (less dominated by large studies) and a limitation (gives more weight to potentially lower-quality small studies).

Q34. A researcher performs 20 independent t-tests on the same dataset comparing treatment vs control on 20 different outcome measures. All are tested at α = 0.05. What is the problem and what is the solution?

A: The problem is multiple comparisons (inflated Type I error). With 20 independent tests at α = 0.05, the probability of at least one false positive is: 1 − (1 − 0.05)20 = 1 − 0.36 = 0.64 (64%). There is a 64% chance of finding at least one "significant" result by chance alone, even if there is no true effect.

Solutions:

  1. Bonferroni correction: Adjust α to 0.05/20 = 0.0025 per test. Simple but very conservative, increases Type II error.
  2. Holm-Bonferroni (step-down): Less conservative, ranks p-values and applies progressively less strict thresholds.
  3. False Discovery Rate (FDR) control (Benjamini-Hochberg): Controls the proportion of false positives among significant results rather than the overall error rate. Less conservative than Bonferroni.
  4. Pre-specify a primary outcome: Designate one primary outcome and treat others as secondary/exploratory, avoids the multiple comparison problem for the primary endpoint.
  5. MANOVA: Use multivariate ANOVA to test all outcomes simultaneously.

Q35. Explain why the STAR*D trial is considered methodologically important in psychiatric research, linking it to concepts of internal and external validity.

A: The STAR*D (Sequenced Treatment Alternatives to Relieve Depression) trial is landmark because it prioritised external validity in a field dominated by explanatory RCTs with strict inclusion criteria.

External validity strengths:

Internal validity trade-offs:

Key findings that shaped clinical practice:

STAR*D exemplifies the efficacy vs effectiveness distinction: it sacrificed some internal validity (no placebo arm, open-label steps) to maximise external validity (generalisable to real patients). This makes its findings more directly applicable to clinical practice than most tightly controlled explanatory RCTs.


Cross-reference: D1 (Study Notes), D2 (Model Answers), D3 (Mnemonics), D4 (Comparisons)

← All study guides