Research September 1, 2026

What an IME actually proves about effort, and where grip testing falls short

When a workers comp or disability claim turns on effort, the examiner gets a hard question: did this person give full effort, or hold back? The tools we lean on are weaker than most reports admit. Here is what grip strength coefficient of variation actually proves, and where it stops.

NK
Nick Kaufman
Flexion Claims · nick@flexionclaims.com
The short answer

Grip strength coefficient of variation is a strong rule-in and a weak rule-out. At the 15 percent cut-off, specificity is 0.94 and sensitivity is 0.55 (Shechtman 2001, PMID 11511013). A high-scatter result is real evidence of submaximal effort. A consistent result clears nobody, because the test misses almost half of the people who hold back. The 0.74 specificity figure repeated in many reviews is wrong: it mixes two operating points from the same paper, and the correct 15 percent value is 0.94. Single-camera video effort scores are a separate problem. Checked against 3D motion capture, one phone camera invents a consistency gap the body never produced, and calibration does not remove it.

When a workers comp or disability claim turns on effort, the examiner gets handed a hard question: did this person give a full effort, or hold back? A lot rides on the answer. Benefits, settlements, and a person's word all sit on it. But the tools we lean on to answer that question are weaker than most reports admit. Here is what one of the most common ones, grip strength coefficient of variation, actually proves, and where it stops.

The grip CV test, and what the numbers really say

The idea is simple. Ask someone to squeeze a dynamometer several times. A person giving full effort tends to produce a tight cluster of readings. Someone holding back tends to scatter. The coefficient of variation, or CV, puts a single number on that scatter. Higher CV, more scatter, more suspicion of submaximal effort.

The best evidence on how well this works comes from Shechtman's 2001 study in the Journal of Hand Therapy PMID 11511013. It reports the area under the ROC curve at 0.79 for three trials and 0.87 for five. At the common 15 percent cut-off, three trials give a sensitivity of 0.55 and a specificity of 0.94.

Sensitivity vs. Specificity at the 15% CV Cut-off
Three-trial grip CV, Shechtman 2001
0.55 Sensitivity misses 45% who hold back 0.94 Specificity flags 94% correctly 0 1.0
Source: Shechtman O., J Hand Ther. 2001;14(3):188-194. PMID 11511013

Read those two numbers slowly, because they point in opposite directions.

Specificity 0.94 means that when the test flags someone as inconsistent, it is usually right. Very few full-effort people scatter that much by chance. So a positive result is a fair rule-in. It is real evidence, not proof, that effort was submaximal.

Sensitivity 0.55 is the other half, and it is the half that gets left out of reports. The test misses almost half of the people who are actually holding back. A clean, consistent result clears nobody. It is a weak rule-out, because plenty of submaximal efforts look perfectly steady on a grip dynamometer.

ROC Curve: Grip CV Diagnostic Accuracy
Area under curve: 0.79 (3 trials), 0.87 (5 trials)
AUC 0.79 (3 trials) 1 - Specificity (false positive rate) Sensitivity (true positive rate) 0 1 1
Source: Shechtman O., J Hand Ther. 2001. AUC improves with trial count.

One more thing, since I chased it down myself. A figure that floats around the reviews, specificity 0.74 at the 15 percent cut-off, is wrong. It comes from mixing two different operating points in the original paper. The real 15 percent specificity is 0.94, which is a harder and more honest bar. A biostatistician flagged the error to me, I bought the full text to settle it, and the tables settle it.

One test is not a verdict

Grip CV earns a place in the file. It does not earn the last word. Used honestly it is a rule-in with teeth and a rule-out with none. An examiner who writes "effort was consistent, therefore the claim is valid" has overread the tool. An examiner who writes "grip CV exceeded 15 percent, consistent with submaximal effort, weighed with the rest of the exam" has it right.

A positive grip CV is worth real weight. A negative one clears nobody. The test is a rule-in with teeth and a rule-out with none, and any report leaning on one number should say how that number was captured.

Where video scoring comes in, and where it breaks

Lately a lot of tools, mine included, have tried to read effort straight from video. Film the movement, track the joints, score the consistency. It is a good idea and I am building toward it. It is also easy to oversell, so I want to be plain about the limits.

A single phone camera has a geometry problem that calibration does not fix. From the side it under-reads how far a joint actually travels. From the front it is worse. Put the same clip against 3D motion-capture ground truth and the range of motion a single camera reports can be off by tens of degrees, with the rep count drifting along with it.

Phone Camera vs. 3D Motion Capture: Consistency Gap
Same movement, same person, two measurement systems
Ground truth p = 0.51 (no gap) Phone camera p = 0.006 (false gap) Movement consistency measurement
Source: REHAB24-6 dataset (CC BY NC), 39 within-subject comparisons, peak speed metric

That matters for one blunt reason: the moment you drop a movement number into a report that a defense expert can pull up and check against a lab, that number has to survive the check, and a lot of them will not.

There is a quieter trap underneath it. On some movements a single camera reports a consistency gap the body never produced. Sloppy form projects into a flat 2D image more erratically than clean form, so the camera shows scatter that the true 3D motion does not. I tested whether a per-setup calibration removes that artifact. It does not. Calibration can fix the level of a number, not the scatter a camera invents.

Calibration Does Not Remove the False Gap
p-values before and after calibration, phone camera vs ground truth
p = 0.006 Before calibration false gap persists p = 0.008 After calibration still significant p = 0.51 Ground truth no real gap
Source: REHAB24-6 dataset (CC BY NC). Calibration fixes the level, not the scatter.

I am saying this out loud because the alternative, a tool that quietly prints a consistency score it cannot defend, is how a whole field loses its credibility in one deposition. A single-camera effort score is not defensible today. A second camera, a depth signal, or a capture gate that throws out the angles that read badly: those are the honest routes. Flexion is taking the capture gate route, and it still has to pass the validation study before any score counts.

What to take from this if you order or rely on effort findings

Three things to carry

1
A positive validity flag is worth real weight. A high grip CV or a failed effort test is a rule-in. It is real evidence, not proof.
2
A negative one clears no one. Consistent output is not proof of full effort. Sensitivity 0.55 means the test misses almost half of submaximal efforts.
3
Any effort number from a single video needs one hard question. Was it checked against ground truth, and from what camera angle? If nobody can answer that, the number is a rough picture, not a measurement.

I am Nick Kaufman. I am building Flexion, a prototype that scores functional-movement effort for comp, disability, and IME work. I would rather it earn trust slowly than claim more than it can prove.

nick@flexionclaims.com

Frequently asked questions

What does a grip strength coefficient of variation test actually prove?

A positive grip CV result (high variability across trials) is real evidence of submaximal effort. At the 15 percent cut-off, specificity is 0.94, meaning very few full-effort people produce that much scatter by chance. A negative result clears nobody, because sensitivity is only 0.55. The test misses almost half of people who are actually holding back.

Can a single phone camera accurately measure effort during an IME?

Not on its own. A single phone camera has a geometry problem that calibration does not fix. Tested against 3D motion-capture ground truth (REHAB24-6), single-camera range of motion can be off by tens of degrees. The phone also reports a false consistency gap on some movements, showing scatter the body never produced (p 0.006 on phone vs p 0.51 in ground truth). Any phone score has to pass a ground-truth check before it goes in a report.

Does calibration fix the single-camera effort scoring problem?

No. Calibration corrects the level of the numbers but not the scatter the camera invents. Every calibration scheme tested still showed the false consistency gap as significant (p 0.008 after calibration). The extra variability is projection noise that does not track the true movement.

What is the correct specificity of grip CV at the 15 percent cut-off?

The correct specificity at the 15 percent cut-off is 0.94, not 0.74. The 0.74 figure that circulates in reviews comes from mixing two different operating points in the original Shechtman 2001 paper. The real 15 percent specificity is 0.94, a harder and more honest bar.