Stereotactic & Functional Neurosurgery

Outcome Instruments in Functional Neurosurgery

Scales, responder thresholds, minimal change, and the difference between improvement and benefit

A score is meaningful only when the instrument fits the construct, the assessment conditions are standardized, and the threshold is interpreted in context. Functional neurosurgery needs absolute change, percent change, adverse effects, and patient-valued function, not a responder label alone.

Evidence status. Thresholds vary by study, baseline severity, intervention, and anchor. This page teaches interpretation and does not reproduce copyrighted scale items. Use licensed, official instruments and current scoring manuals. Scale ranges, subscale names, and scoring structure are facts and are stated here; item wording and printed forms are the copyrighted part and are not.

Orientation

Outcome instruments serve at least four jobs: describe baseline severity, quantify change, define trial eligibility or response, and compare groups. One scale may not do all four well. A statistically significant mean change can be clinically trivial; a responder threshold can conceal important partial improvement or harm.

Prespecify the instrument, assessment state, time point, calculation, missing-data rule, and responder definition before treatment. Then pair disease scores with function, quality of life, and adverse effects.

Part I

Interpretation before arithmetic

1.Absolute, percent, and clinically important change

Absolute change preserves scale units. Percent change is intuitive but unstable near low baseline values. A minimal clinically important difference is an anchor- or distribution-based estimate of meaningful change, not a universal biological cutoff. Response and remission answer different questions.

Always report baseline and follow-up distributions beside a responder percentage. A patient can cross a threshold by one point; another can improve substantially and remain just outside it.

2.State and rater matter

Parkinson motor scores depend on medication and stimulation state. Dystonia ratings depend on video angle, task, duration, and whether disability or pain is included. Depression and OCD scales depend on trained interviewing and time frame. Trigeminal neuralgia scores may disagree with patient-reported pain interference.

Measurement ruleWrite the assessment condition as part of the outcome: medication OFF/stimulation ON, blinded video rating, independent psychiatric rater, or patient-reported seven-day window. A score without its state is incomplete data.
Part II

Common instruments

InstrumentConstructRange and structureCommon interpretationImportant caution
MDS-UPDRSParkinson nonmotor, daily living, motor examination, complicationsPart I 0–52, Part II 0–52, Part III 0–132, Part IV 0–24; total 0–260Part III motor change is often central to DBS studiesPermission and rater training required; specify medication and stimulation state
Y-BOCSOCD symptom severity0–40; obsession and compulsion subtotals 0–20 eachConsensus response: at least 35% plus CGI-I 1–2; remission uses diagnostic status or a score ≤12 plus CGI-S 1–2, with duration criteriaSelf-report and interview versions are not interchangeable
MADRSDepressive symptom severity0–60; ten items scored 0–6Trial convention: response at least 50%; remission 10 or lowerThresholds have empirical support but vary by definition; assess suicidality outside the total alone
TWSTRSCervical dystonia severity, disability, pain0–85 total, with severity, disability, and pain subscalesReport total and subscales; minimal important change about 8 points on the totalThe 8-point figure comes from botulinum toxin practice, not from a DBS cohort
BFMDRSGeneralized and segmental dystoniaMovement and disability subscales, scored separatelyThe currency of the published dystonia DBS literatureRecord movement and disability at baseline; they diverge as deformity becomes fixed
Pain numeric rating scaleChronic pain intensity0–10Responder conventionally at 50% reduction, with 30% as a lower tierConventions from analgesic drug trials, not thresholds derived in implanted patients
Seizure diarySeizure frequencyA count, not a scoreResponder conventionally at 50% or greater reductionPatient-kept and substantially incomplete; corroborate where the question will bear it
BNI pain scoreTrigeminal neuralgia pain and medication categoryOrdinal grades I, II, IIIa, IIIb, IV, VGrade I is pain-free off medication; grades I–IIIb are adequate controlRead which band a paper reports before you read its percentage

3.MDS-UPDRS

The Movement Disorder Society revision has four parts and requires permission for use and standardized rater training. Part I covers nonmotor experiences of daily living and runs 0 to 52, Part II motor experiences of daily living 0 to 52, Part III the motor examination 0 to 132, and Part IV motor complications 0 to 24, for a total of 0 to 260. Carry those denominators: a six-point change means one thing on a 132-point motor examination and something else on a 24-point complications score. DBS studies commonly emphasize Part III but should also include Part II daily living, motor complications, cognition, gait/falls, medication dose, and adverse effects. Specify the four medication/stimulation combinations when the scientific question requires them. Anchor-based minimal-change estimates for Part III exist in the Parkinson literature; take the value from the primary source you intend to cite rather than from any secondary summary, including this one.

4.Y-BOCS and MADRS

For OCD, the international consensus combines a Y-BOCS reduction of at least 35% with CGI-I 1 or 2 for at least one week to define response. Remission means no longer meeting diagnostic criteria for at least one week; when a structured interview is unavailable, the operational alternative is Y-BOCS 12 or lower plus CGI-S 1 or 2 for at least one week. A Y-BOCS-only endpoint should be identified as such. The empirical support is close but not identical. Signal-detection work in 288 patients pooled from four randomized trials found that a reduction of at least 35% best predicted response, that a posttreatment score of 14 or lower best predicted response by global impression, and that 12 or lower best predicted wellness (remission together with good quality of life and adaptive functioning), a stricter construct than remission alone. A 2024 individual-patient meta-analysis of 25 randomized trials and 1,235 participants found the optimal empirical thresholds were lower, at 30% for response and 15 for remission, but recommended keeping the consensus definitions because contemporary trials enroll more refractory patients. Prespecify the consensus values, report continuous change beside them, and state which definition you used: moving the line from 35% to 30% moves patients across it.

For depression, trials conventionally use a 50% MADRS reduction for response and a score of 10 or lower for remission. These are common conventions with empirical validation literature, rather than universal cutoffs. Studies anchored to global severity or diagnostic remission have produced different MADRS thresholds; state whether the protocol uses less than 10 or 10 or lower, since they are not identical. These conventions facilitate comparison; they do not substitute for function, durability, or safety.

5.TWSTRS, BFMDRS, and BNI

TWSTRS runs to 85 points and separates severity, disability, and pain. Report the components, because the same total can arise from different clinical changes. Clinically important change on TWSTRS is anchored, not intrinsic. The most-cited estimate comes from CD PROBE, a 479-patient real-world registry of onabotulinumtoxinA for cervical dystonia: against the patient global impression of change, total-score improvements of 8, 9, and 11 points corresponded to minimally improved, much improved, and very much improved, and the authors settled on 8 points as the minimal clinically important change. Two cautions before carrying that number into a DBS clinic. It was derived under chemodenervation, in patients whose mean baseline total was 39, and an anchor-based threshold does not travel cleanly to a different intervention, a different baseline severity, or a different follow-up interval. CD PROBE also found the total mapped to global impression better than any single subscale, which is a reason to report the total, not a reason to stop reporting severity, disability, and pain separately, since the same total change can be all pain or all severity.

TWSTRS itself has been revised. The Comprehensive Cervical Dystonia Rating Scale packages a modified TWSTRS-2, with subscales for motor severity, disability, and pain, alongside a psychiatric screening module and a disease-specific quality-of-life measure, and it has been clinimetrically tested. If you are designing a cervical dystonia protocol now, choose deliberately between classic TWSTRS, which is what the historical DBS literature is expressed in, and the comprehensive scale, which measures more of what the patient notices.

For generalized and segmental dystonia the instrument is the Burke-Fahn-Marsden Dystonia Rating Scale, and nearly every published DBS number you will meet is a BFMDRS movement percentage. It carries a movement subscale and a separate disability subscale. Record both at baseline and report both: they diverge as fixed deformity accumulates, and a cohort whose movement score falls by two-thirds while disability barely moves has told you something about those patients that the movement score alone conceals.

The BNI pain intensity score is efficient for trigeminal neuralgia because it folds medication into the category: I, no pain and no medication; II, occasional pain not requiring medication; IIIa, no pain on medication; IIIb, pain adequately controlled on medication; IV, pain not adequately controlled; V, no relief. Two success definitions circulate and they are not the same number. Grade I is pain-free off medication. Grades I through IIIb mean adequate control with or without medication, and that is the band most radiosurgical and percutaneous series report. Read the definition before you read the percentage. The scale remains ordinal, so add attack frequency, triggers, complications, and patient-reported quality of life; facial numbness has its own separate BNI score and belongs in its own column rather than folded into the pain grade.

Permissions before protocolSettle instrument permissions before the protocol is written, not after the first patient. MDS-UPDRS requires a permission request through the Movement Disorder Society, rater training for the motor examination, and the official current form in the approved translation for your site. For every other instrument on this page, confirm the current rights holder and the terms in writing before you print a form, build a database, or reproduce anything in a manuscript or a talk. Verify the applicable license before reproducing items or forms; permissions and exceptions depend on the instrument, version, and proposed use.

6.Pain, seizures, tremor, and spasticity

Pain. Intensity on a 0 to 10 numeric rating scale is the near-universal primary, and the near-universal responder definition is a 50% reduction, with 30% widely treated as a smaller but still meaningful tier. Both are conventions carried over from analgesic drug trials rather than thresholds derived in implanted patients, and a single 50% line can make a device therapy's success rate look cleaner than the underlying distribution. Intensity alone is insufficient: add an interference or disability measure, a global impression of change, opioid dose, and the number most often missing, which is explantation and device revision reported on the same denominator as the responder rate. A cohort with 70% responders and a quarter of devices explanted at five years has not been described by its responder rate.

Seizures. The instrument behind essentially every neuromodulation figure in epilepsy is a patient-kept diary, and the convention is a 50% or greater reduction in frequency from a prospectively recorded baseline. Seizure diaries have important ascertainment limitations. In video-EEG monitoring of 91 adults with focal epilepsy, patients failed to document 55.5% of recorded seizures, and the omissions largely reflected seizure unawareness (including seizures arising from sleep) rather than poor compliance, so reminders alone do not fix the problem. A diary-based reduction estimates reported seizures; its relationship to total clinical and electrographic seizure burden depends on stable ascertainment and the seizure types being counted. Require an adequate prospective baseline, define a countable event before the study starts, and corroborate the diary electrographically where the question will bear it.

Tremor and spasticity. Tremor outcomes commonly use the Fahn-Tolosa-Marin Clinical Rating Scale for Tremor (FTM-CRST) or the Essential Tremor Rating Assessment Scale (TETRAS), with accelerometry and spiral ratings as adjuncts; specify which limb, which task, and whether the rater was blinded. Spasticity and rhizotomy outcomes need a function measure and a tone measure together: GMFCS stratification, GMFM for gross motor function, and modified Ashworth or Tardieu for tone, with instrumented gait when it will change a decision. In both, accompany group estimates with uncertainty, individual trajectories, function, and adverse outcomes.

Part III

Build an outcome architecture

7.Core, disease, and mechanism layers

Use a core layer across procedures: serious adverse events, reoperation, infection, cognition, mood, quality of life, and patient goal attainment. Add a disease layer such as MDS-UPDRS, Y-BOCS, MADRS, TWSTRS, BFMDRS, a pain numeric rating scale, a seizure diary, or the BNI grade. Add a mechanism layer matched to the intervention: tremor accelerometry, LFP burden, pain unpleasantness, gait sensors, or stimulation energy.

8.Responder analyses without threshold theater

Prespecify one primary threshold and provide sensitivity analyses at adjacent thresholds. Report continuous change and confidence intervals. Show deterioration as well as improvement. When several scales are tested, identify the hierarchy and correct or contextualize multiplicity.

Do not manufacture responseAvoid choosing the cutoff after seeing the data, using the best of several follow-up visits, or excluding explants and nonresponders from the denominator. Define how deaths, revisions, and missing assessments are handled.

9.Time is an outcome

Early microlesion effects, programming optimization, placebo, disease progression, tolerance, and hardware changes alter trajectories. Choose time points that match the intervention and include durability. A useful report shows each patient's course, not just group means at a convenient visit.

Two design issues that change interpretationBlinding and comparator choice change interpretation. Open-label improvement may include expectancy, cointerventions, regression to the mean, and observer effects; it is not interchangeable with a randomized between-group treatment effect. This does not imply that every blinded estimate must be smaller. Long-term results can be affected by attrition, but an increasing responder rate is not proof of selective retention. Report the number contributing at each time point, reasons for missingness, the prespecified estimand, and sensitivity analyses. Report explants, revisions, and deaths on an explicit denominator; consider both observed-case and conservative analyses without assuming every missing observation is treatment failure.
Board and clinic pearls
  • A responder threshold is a convention, not a natural law.
  • Carry the denominator: a point change means nothing without the scale range.
  • Always specify medication, stimulation, rater, and timing conditions.
  • Do not reproduce proprietary or copyrighted scale content without permission.
  • Report continuous change and adverse outcomes beside responder rates.
  • Grade I and grades I–IIIb are different BNI outcomes; check which one a paper reports.
  • Interpret blinding, comparators, retention, and missing-data assumptions before comparing response rates.
  • Match one outcome layer to disease and another to the mechanism being claimed.
Sources

Selected References

Selected for trainees. Asterisked entries are the best starting points.

  1. Movement Disorder Society. MDS-UPDRS: official scale information, training, and permissions. Permission requests, rater training, official forms and approved translations. MDS
  2. Farris SG, et al. Treatment response, symptom remission, and wellness in obsessive-compulsive disorder. J Clin Psychiatry. 2013;74(7):685–690. Signal-detection analysis of 288 patients pooled from four randomized trials; source of the 35% response threshold and of the 12-or-lower wellness cutoff. PubMed
  3. Ramakrishnan D, et al. An evaluation of treatment response and remission definitions in adult obsessive-compulsive disorder: a systematic review and individual-patient data meta-analysis. J Psychiatr Res. 2024;173:387–397. Individual-patient-data meta-analysis, 25 randomized trials, 1,235 participants; optimal empirical thresholds of 30% and 15, with the consensus definitions nonetheless recommended. PubMed
  4. Montgomery SA, Asberg M. A new depression scale designed to be sensitive to change. Br J Psychiatry. 1979;134(4):382–389. The original MADRS. It is the source for the instrument, not for the 50% and 10-or-lower conventions. PubMed
  5. Dashtipour K, et al. Minimal clinically important change in patients with cervical dystonia: results from the CD PROBE study. J Neurol Sci. 2019;405:116413. Prospective real-world registry of onabotulinumtoxinA in 479 patients; anchor-based minimal clinically important change against patient global impression, not a DBS cohort. PubMed
  6. Comella CL, Perlmutter JS, Jinnah HA, et al. Clinimetric testing of the comprehensive cervical dystonia rating scale. Mov Disord. 2016;31(4):563–569. The modular revision: TWSTRS-2 severity, disability, and pain, with psychiatric screening and quality of life. PubMed
  7. Hoppe C, Poepel A, Elger CE. Epilepsy: accuracy of patient seizure counts. Arch Neurol. 2007;64(11):1595–1599. Ninety-one adults on video-EEG failed to document 55.5% of recorded seizures; omissions largely reflected seizure unawareness rather than poor compliance. PubMed
  8. Venda Nova C, et al. Treatment outcomes in trigeminal neuralgia: a systematic review of domains, dimensions and measures. World Neurosurg X. 2020;6:100070. Systematic review of outcome domains, dimensions, and measures in trigeminal neuralgia; the case for reporting beyond an ordinal pain grade. PubMed
  9. Mataix-Cols D, et al. Towards an international expert consensus for defining treatment response, remission, recovery and relapse in obsessive-compulsive disorder. World Psychiatry. 2016;15(1):80–81. PubMed
  10. Hawley CJ, et al. Defining remission by cut off score on the MADRS: selecting the optimal value. J Affect Disord. 2002;72(2):177–184. PubMed
  11. Zimmerman M, et al. Defining remission on the Montgomery-Asberg depression rating scale. J Clin Psychiatry. 2004;65(2):163–168. PubMed
  12. Ondo W, et al. Comparison of the Fahn-Tolosa-Marin Clinical Rating Scale and the Essential Tremor Rating Assessment Scale. Mov Disord Clin Pract. 2018;5(1):60–65. PubMed