Stereotactic & Functional Neurosurgery
Outcome Instruments in Functional Neurosurgery
Scales, responder thresholds, minimal change, and the difference between improvement and benefit
A score is meaningful only when the instrument fits the construct, the assessment conditions are standardized, and the threshold is interpreted in context. Functional neurosurgery needs absolute change, percent change, adverse effects, and patient-valued function, not a responder label alone.
Evidence status. Thresholds vary by study, baseline severity, intervention, and anchor. This page teaches interpretation and does not reproduce copyrighted scale items. Use licensed, official instruments and current scoring manuals. Scale ranges, subscale names, and scoring structure are facts and are stated here; item wording and printed forms are the copyrighted part and are not.
Orientation
Outcome instruments serve at least four jobs: describe baseline severity, quantify change, define trial eligibility or response, and compare groups. One scale may not do all four well. A statistically significant mean change can be clinically trivial; a responder threshold can conceal important partial improvement or harm.
Prespecify the instrument, assessment state, time point, calculation, missing-data rule, and responder definition before treatment. Then pair disease scores with function, quality of life, and adverse effects.
Interpretation before arithmetic
1.Absolute, percent, and clinically important change
Absolute change preserves scale units. Percent change is intuitive but unstable near low baseline values. A minimal clinically important difference is an anchor- or distribution-based estimate of meaningful change, not a universal biological cutoff. Response and remission answer different questions.
Always report baseline and follow-up distributions beside a responder percentage. A patient can cross a threshold by one point; another can improve substantially and remain just outside it.
2.State and rater matter
Parkinson motor scores depend on medication and stimulation state. Dystonia ratings depend on video angle, task, duration, and whether disability or pain is included. Depression and OCD scales depend on trained interviewing and time frame. Trigeminal neuralgia scores may disagree with patient-reported pain interference.
Common instruments
| Instrument | Construct | Range and structure | Common interpretation | Important caution |
|---|---|---|---|---|
| MDS-UPDRS | Parkinson nonmotor, daily living, motor examination, complications | Part I 0–52, Part II 0–52, Part III 0–132, Part IV 0–24; total 0–260 | Part III motor change is often central to DBS studies | Permission and rater training required; specify medication and stimulation state |
| Y-BOCS | OCD symptom severity | 0–40; obsession and compulsion subtotals 0–20 each | Consensus response: at least 35% plus CGI-I 1–2; remission uses diagnostic status or a score ≤12 plus CGI-S 1–2, with duration criteria | Self-report and interview versions are not interchangeable |
| MADRS | Depressive symptom severity | 0–60; ten items scored 0–6 | Trial convention: response at least 50%; remission 10 or lower | Thresholds have empirical support but vary by definition; assess suicidality outside the total alone |
| TWSTRS | Cervical dystonia severity, disability, pain | 0–85 total, with severity, disability, and pain subscales | Report total and subscales; minimal important change about 8 points on the total | The 8-point figure comes from botulinum toxin practice, not from a DBS cohort |
| BFMDRS | Generalized and segmental dystonia | Movement and disability subscales, scored separately | The currency of the published dystonia DBS literature | Record movement and disability at baseline; they diverge as deformity becomes fixed |
| Pain numeric rating scale | Chronic pain intensity | 0–10 | Responder conventionally at 50% reduction, with 30% as a lower tier | Conventions from analgesic drug trials, not thresholds derived in implanted patients |
| Seizure diary | Seizure frequency | A count, not a score | Responder conventionally at 50% or greater reduction | Patient-kept and substantially incomplete; corroborate where the question will bear it |
| BNI pain score | Trigeminal neuralgia pain and medication category | Ordinal grades I, II, IIIa, IIIb, IV, V | Grade I is pain-free off medication; grades I–IIIb are adequate control | Read which band a paper reports before you read its percentage |
3.MDS-UPDRS
The Movement Disorder Society revision has four parts and requires permission for use and standardized rater training. Part I covers nonmotor experiences of daily living and runs 0 to 52, Part II motor experiences of daily living 0 to 52, Part III the motor examination 0 to 132, and Part IV motor complications 0 to 24, for a total of 0 to 260. Carry those denominators: a six-point change means one thing on a 132-point motor examination and something else on a 24-point complications score. DBS studies commonly emphasize Part III but should also include Part II daily living, motor complications, cognition, gait/falls, medication dose, and adverse effects. Specify the four medication/stimulation combinations when the scientific question requires them. Anchor-based minimal-change estimates for Part III exist in the Parkinson literature; take the value from the primary source you intend to cite rather than from any secondary summary, including this one.
4.Y-BOCS and MADRS
For OCD, the international consensus combines a Y-BOCS reduction of at least 35% with CGI-I 1 or 2 for at least one week to define response. Remission means no longer meeting diagnostic criteria for at least one week; when a structured interview is unavailable, the operational alternative is Y-BOCS 12 or lower plus CGI-S 1 or 2 for at least one week. A Y-BOCS-only endpoint should be identified as such. The empirical support is close but not identical. Signal-detection work in 288 patients pooled from four randomized trials found that a reduction of at least 35% best predicted response, that a posttreatment score of 14 or lower best predicted response by global impression, and that 12 or lower best predicted wellness (remission together with good quality of life and adaptive functioning), a stricter construct than remission alone. A 2024 individual-patient meta-analysis of 25 randomized trials and 1,235 participants found the optimal empirical thresholds were lower, at 30% for response and 15 for remission, but recommended keeping the consensus definitions because contemporary trials enroll more refractory patients. Prespecify the consensus values, report continuous change beside them, and state which definition you used: moving the line from 35% to 30% moves patients across it.
For depression, trials conventionally use a 50% MADRS reduction for response and a score of 10 or lower for remission. These are common conventions with empirical validation literature, rather than universal cutoffs. Studies anchored to global severity or diagnostic remission have produced different MADRS thresholds; state whether the protocol uses less than 10 or 10 or lower, since they are not identical. These conventions facilitate comparison; they do not substitute for function, durability, or safety.
5.TWSTRS, BFMDRS, and BNI
TWSTRS runs to 85 points and separates severity, disability, and pain. Report the components, because the same total can arise from different clinical changes. Clinically important change on TWSTRS is anchored, not intrinsic. The most-cited estimate comes from CD PROBE, a 479-patient real-world registry of onabotulinumtoxinA for cervical dystonia: against the patient global impression of change, total-score improvements of 8, 9, and 11 points corresponded to minimally improved, much improved, and very much improved, and the authors settled on 8 points as the minimal clinically important change. Two cautions before carrying that number into a DBS clinic. It was derived under chemodenervation, in patients whose mean baseline total was 39, and an anchor-based threshold does not travel cleanly to a different intervention, a different baseline severity, or a different follow-up interval. CD PROBE also found the total mapped to global impression better than any single subscale, which is a reason to report the total, not a reason to stop reporting severity, disability, and pain separately, since the same total change can be all pain or all severity.
TWSTRS itself has been revised. The Comprehensive Cervical Dystonia Rating Scale packages a modified TWSTRS-2, with subscales for motor severity, disability, and pain, alongside a psychiatric screening module and a disease-specific quality-of-life measure, and it has been clinimetrically tested. If you are designing a cervical dystonia protocol now, choose deliberately between classic TWSTRS, which is what the historical DBS literature is expressed in, and the comprehensive scale, which measures more of what the patient notices.
For generalized and segmental dystonia the instrument is the Burke-Fahn-Marsden Dystonia Rating Scale, and nearly every published DBS number you will meet is a BFMDRS movement percentage. It carries a movement subscale and a separate disability subscale. Record both at baseline and report both: they diverge as fixed deformity accumulates, and a cohort whose movement score falls by two-thirds while disability barely moves has told you something about those patients that the movement score alone conceals.
The BNI pain intensity score is efficient for trigeminal neuralgia because it folds medication into the category: I, no pain and no medication; II, occasional pain not requiring medication; IIIa, no pain on medication; IIIb, pain adequately controlled on medication; IV, pain not adequately controlled; V, no relief. Two success definitions circulate and they are not the same number. Grade I is pain-free off medication. Grades I through IIIb mean adequate control with or without medication, and that is the band most radiosurgical and percutaneous series report. Read the definition before you read the percentage. The scale remains ordinal, so add attack frequency, triggers, complications, and patient-reported quality of life; facial numbness has its own separate BNI score and belongs in its own column rather than folded into the pain grade.
6.Pain, seizures, tremor, and spasticity
Pain. Intensity on a 0 to 10 numeric rating scale is the near-universal primary, and the near-universal responder definition is a 50% reduction, with 30% widely treated as a smaller but still meaningful tier. Both are conventions carried over from analgesic drug trials rather than thresholds derived in implanted patients, and a single 50% line can make a device therapy's success rate look cleaner than the underlying distribution. Intensity alone is insufficient: add an interference or disability measure, a global impression of change, opioid dose, and the number most often missing, which is explantation and device revision reported on the same denominator as the responder rate. A cohort with 70% responders and a quarter of devices explanted at five years has not been described by its responder rate.
Seizures. The instrument behind essentially every neuromodulation figure in epilepsy is a patient-kept diary, and the convention is a 50% or greater reduction in frequency from a prospectively recorded baseline. Seizure diaries have important ascertainment limitations. In video-EEG monitoring of 91 adults with focal epilepsy, patients failed to document 55.5% of recorded seizures, and the omissions largely reflected seizure unawareness (including seizures arising from sleep) rather than poor compliance, so reminders alone do not fix the problem. A diary-based reduction estimates reported seizures; its relationship to total clinical and electrographic seizure burden depends on stable ascertainment and the seizure types being counted. Require an adequate prospective baseline, define a countable event before the study starts, and corroborate the diary electrographically where the question will bear it.
Tremor and spasticity. Tremor outcomes commonly use the Fahn-Tolosa-Marin Clinical Rating Scale for Tremor (FTM-CRST) or the Essential Tremor Rating Assessment Scale (TETRAS), with accelerometry and spiral ratings as adjuncts; specify which limb, which task, and whether the rater was blinded. Spasticity and rhizotomy outcomes need a function measure and a tone measure together: GMFCS stratification, GMFM for gross motor function, and modified Ashworth or Tardieu for tone, with instrumented gait when it will change a decision. In both, accompany group estimates with uncertainty, individual trajectories, function, and adverse outcomes.
Build an outcome architecture
7.Core, disease, and mechanism layers
Use a core layer across procedures: serious adverse events, reoperation, infection, cognition, mood, quality of life, and patient goal attainment. Add a disease layer such as MDS-UPDRS, Y-BOCS, MADRS, TWSTRS, BFMDRS, a pain numeric rating scale, a seizure diary, or the BNI grade. Add a mechanism layer matched to the intervention: tremor accelerometry, LFP burden, pain unpleasantness, gait sensors, or stimulation energy.
8.Responder analyses without threshold theater
Prespecify one primary threshold and provide sensitivity analyses at adjacent thresholds. Report continuous change and confidence intervals. Show deterioration as well as improvement. When several scales are tested, identify the hierarchy and correct or contextualize multiplicity.
9.Time is an outcome
Early microlesion effects, programming optimization, placebo, disease progression, tolerance, and hardware changes alter trajectories. Choose time points that match the intervention and include durability. A useful report shows each patient's course, not just group means at a convenient visit.
- A responder threshold is a convention, not a natural law.
- Carry the denominator: a point change means nothing without the scale range.
- Always specify medication, stimulation, rater, and timing conditions.
- Do not reproduce proprietary or copyrighted scale content without permission.
- Report continuous change and adverse outcomes beside responder rates.
- Grade I and grades I–IIIb are different BNI outcomes; check which one a paper reports.
- Interpret blinding, comparators, retention, and missing-data assumptions before comparing response rates.
- Match one outcome layer to disease and another to the mechanism being claimed.
Selected References
Selected for trainees. Asterisked entries are the best starting points.
- Movement Disorder Society. MDS-UPDRS: official scale information, training, and permissions. Permission requests, rater training, official forms and approved translations. MDS
- Farris SG, et al. Treatment response, symptom remission, and wellness in obsessive-compulsive disorder. J Clin Psychiatry. 2013;74(7):685–690. Signal-detection analysis of 288 patients pooled from four randomized trials; source of the 35% response threshold and of the 12-or-lower wellness cutoff. PubMed
- Ramakrishnan D, et al. An evaluation of treatment response and remission definitions in adult obsessive-compulsive disorder: a systematic review and individual-patient data meta-analysis. J Psychiatr Res. 2024;173:387–397. Individual-patient-data meta-analysis, 25 randomized trials, 1,235 participants; optimal empirical thresholds of 30% and 15, with the consensus definitions nonetheless recommended. PubMed
- Montgomery SA, Asberg M. A new depression scale designed to be sensitive to change. Br J Psychiatry. 1979;134(4):382–389. The original MADRS. It is the source for the instrument, not for the 50% and 10-or-lower conventions. PubMed
- Dashtipour K, et al. Minimal clinically important change in patients with cervical dystonia: results from the CD PROBE study. J Neurol Sci. 2019;405:116413. Prospective real-world registry of onabotulinumtoxinA in 479 patients; anchor-based minimal clinically important change against patient global impression, not a DBS cohort. PubMed
- Comella CL, Perlmutter JS, Jinnah HA, et al. Clinimetric testing of the comprehensive cervical dystonia rating scale. Mov Disord. 2016;31(4):563–569. The modular revision: TWSTRS-2 severity, disability, and pain, with psychiatric screening and quality of life. PubMed
- Hoppe C, Poepel A, Elger CE. Epilepsy: accuracy of patient seizure counts. Arch Neurol. 2007;64(11):1595–1599. Ninety-one adults on video-EEG failed to document 55.5% of recorded seizures; omissions largely reflected seizure unawareness rather than poor compliance. PubMed
- Venda Nova C, et al. Treatment outcomes in trigeminal neuralgia: a systematic review of domains, dimensions and measures. World Neurosurg X. 2020;6:100070. Systematic review of outcome domains, dimensions, and measures in trigeminal neuralgia; the case for reporting beyond an ordinal pain grade. PubMed
- Mataix-Cols D, et al. Towards an international expert consensus for defining treatment response, remission, recovery and relapse in obsessive-compulsive disorder. World Psychiatry. 2016;15(1):80–81. PubMed
- Hawley CJ, et al. Defining remission by cut off score on the MADRS: selecting the optimal value. J Affect Disord. 2002;72(2):177–184. PubMed
- Zimmerman M, et al. Defining remission on the Montgomery-Asberg depression rating scale. J Clin Psychiatry. 2004;65(2):163–168. PubMed
- Ondo W, et al. Comparison of the Fahn-Tolosa-Marin Clinical Rating Scale and the Essential Tremor Rating Assessment Scale. Mov Disord Clin Pract. 2018;5(1):60–65. PubMed