EVIDENCE-BASED CLINICAL EXAMINATION  —  CHAPTER 2

The Chest Examination

From Practice to Evidence · What do two physicians agree on when they examine the same chest?

Kappa 0.85The ERS nomenclature of 2016Lung ultrasound

A sign that two competent examiners cannot agree on is not a measurement. It may still be a ritual, a habit, or a courtesy to the patient, but it cannot carry the weight of a diagnosis, because no one can say which of the two examiners the sign belongs to. That is the test the “rational clinical examination” literature has applied to the bedside since the JAMA series of 1992: before asking how accurate a sign is, ask whether it is reproducible, since a sign that is not reproducible cannot be accurate except by luck.

The chest is where this question bites hardest. It is the examination students learn first and most completely, it has the richest vocabulary of signs, and it has been studied for reproducibility more than any other region. What follows is the examination as it should be performed, and then what the studies say about each part of it.

The examination as it should be performed

Setting up. The patient sits at forty-five degrees with the chest exposed, in good light, and is examined from the right and from the foot of the bed. Before touching the chest, spend thirty seconds watching: the effort of breathing, the position the patient has chosen, whether they can speak in full sentences, any audible sound such as stridor or wheeze, and what stands at the bedside, the oxygen, the inhalers, the sputum pot.

The peripheral survey.

Inspection of the chest.

Palpation.

Percussion. Compare like with like, moving from side to side: the clavicles first, then the anterior chest, the axillae, and the posterior chest with the scapulae drawn apart. The notes are resonant in health, hyper-resonant in pneumothorax and emphysema, dull in consolidation, collapse and fibrosis, and stony dull in effusion. Percuss out the liver and the cardiac dullness; their loss suggests hyperinflation. Auscultatory percussion, the stethoscope on the posterior chest while the sternum is tapped, has been advocated for effusion.

Auscultation. With the diaphragm, the patient breathing through an open mouth, at least six sites on each side compared symmetrically, the axillae and apices included. Record the breath sounds by intensity (normal, reduced, absent) and character, vesicular or bronchial, where bronchial breathing has an expiratory phase as long as the inspiratory one with a pause between. Record the added sounds in the nomenclature of the European Respiratory Society of 2016: crackles, fine and late-inspiratory in fibrosis and oedema, coarse and early or mid-inspiratory in bronchiectasis and retained secretions, where they clear with a cough; wheezes, high-pitched and continuous, mainly expiratory, polyphonic in asthma and COPD, monophonic and fixed over a single obstructed bronchus; rhonchi, the low-pitched wheeze of secretions; stridor, inspiratory, from the upper airway; and the pleural rub. Then vocal resonance, whispering pectoriloquy, and aegophony, the change of E to A over the upper border of an effusion or over consolidation.

Completing the examination. Sacral and ankle oedema, the temperature chart, the oxygen saturation, a peak flow or spirometry, the sputum, and a request for the chest radiograph. Then the findings are put together into a pattern, consolidation, effusion, pneumothorax, collapse, fibrosis or hyperinflation, rather than left as a list of signs.

When two physicians examine one chest

The classic studies of the sixties, seventies and eighties put several physicians in front of the same patients and counted how often they agreed. Smyllie in 1965, Gjørup in 1984 and Spiteri in 1988 all reached the same conclusion: on most signs, agreement was poor. The measure they used, and the one still used, is kappa, which counts agreement beyond what chance alone would produce. A kappa of one is perfect agreement, zero is what two examiners would reach by tossing coins, and values below about 0.4 are usually called poor.

In the Danish study, two hundred and two unselected medical inpatients were examined for eleven signs by pairs of physicians, with and without knowledge of the history. The kappa values for the signs ranged from −0.04 to 0.58, and for the diagnoses from 0.15 to 0.68, and knowing the history did not make the examiners agree more. In Spiteri’s series, twenty-four physicians in sets of four examined patients with well-defined chest signs, and all four agreed only fifty-five per cent of the time. A later review put it flatly: interobserver agreement on respiratory signs has been studied repeatedly and generally found to be low, and so has the correlation between the signs and lung function.

The picture is more useful sign by sign. The figures below are approximate, pooled from Spiteri, Wipf and McGee’s Evidence-Based Physical Diagnosis.

SignKappaComment
Asymmetric chest expansionabout 0.85The most reproducible sign. In McGee’s figures it carries a positive likelihood ratio of 44 for pneumonia, very specific, but with a sensitivity of about 4 per cent.
Wheezesabout 0.5Consistently among the better signs.
Crackles0.3 to 0.6Better in the lateral decubitus position, where agreement reached about 0.5.
Reduced breath soundsabout 0.4Moderate.
Percussion dullnessabout 0.5Moderate, and best for effusion.
Bronchial breathingabout 0.3Fair.
Tactile fremitus0.0 to 0.3Near chance in Spiteri’s series. A later intensive-care study of palpable chest-wall fremitus in ventilated patients found agreement depends on the zone: moderate to almost perfect in the upper zones, less than chance to moderate in the lower ones.
Whispering pectoriloquyabout 0.1Poor.
Clubbingabout 0.45Early clubbing is hard to be sure of; objective indices such as the phalangeal depth ratio have been proposed.

Reproducibility is not accuracy, and the accuracy figures are no kinder. In Wipf’s study three examiners’ clinical diagnosis of pneumonia had a sensitivity of 47 to 69 per cent and a specificity of 58 to 75 per cent, and the authors concluded that the traditional chest examination is not accurate enough on its own to confirm or exclude pneumonia. The literature on pleural effusion is more favourable. Kalantri’s study of 278 patients in rural India, with two examiners blind to the radiograph and to each other, found positive likelihood ratios between 1.5 and 8.1 for the individual signs, and Wong’s review in the JAMA series of 2009 found dullness to percussion and reduced tactile fremitus to be the useful signs, while auscultation added little.

What has changed since 2020

Lung ultrasound has become the comparator, and the examination loses. A systematic review of mechanically ventilated intensive-care patients framed its question around the poor diagnostic accuracy of chest radiography and auscultation, and the need for something more reliable. In the emergency department, Zanobetti’s study of 2,683 patients with undifferentiated dyspnoea found that point-of-care ultrasound cut the time to diagnosis from 186 minutes to 24, performed as well as the standard work-up of examination plus chest radiograph, and was more sensitive for heart failure, with an overall concordance of kappa 0.71. For pleural effusion, a study of 34 inpatients in Calgary and Spokane, one of the few to compare ultrasound with a full physical examination rather than with the stethoscope alone, found the examination detected only 44 per cent of effusions against 98 per cent for ultrasound. Lung ultrasound has reached the ambulatory setting too: a 2024 review of six paediatric studies found a pooled sensitivity of 91 per cent for community-acquired pneumonia, and an updated meta-analysis in adults in 2025 found it more sensitive than the chest radiograph.

Digital and machine-assisted auscultation is the second front. Its motive is precisely the reproducibility problem: the ordinary stethoscope carries inter-listener variability and subjectivity, and cannot record a sound for later review or for telemedicine. Park’s meta-analysis of 2025 on paediatric lung-sound classifiers found pooled sensitivity and specificity above 90 per cent for wheeze, and concluded that machine-learning models reach high accuracy but are limited by heterogeneous datasets, the absence of standard guidelines, and little external validation. In the same year Cox and the FEND-TB consortium tested a commercial device against a microbiological reference standard for tuberculosis, enrolling 240 symptomatic adults in South Africa, Uganda, Vietnam and Peru, with sounds from six auscultation positions analysed blind by the manufacturer. The direction of travel is screening in low-resource settings, not the replacement of bedside diagnosis.

Human listening remains limited. Kim and colleagues in Korea played recorded lung sounds to seventy trainees, from medical students to fellows, and measured correct response rates of 73.5 per cent for normal sounds, 72.2 per cent for crackles, 56.3 per cent for wheezes and 41.7 per cent for rhonchi, which accounts for much of the kappa data above. Training helps, and the point is not confined to the chest: in a structured examination of the upper airway in patients with sleep-disordered breathing, examiners with a year’s training in sleep medicine agreed at kappa 0.69, which counts as good, against 0.48 for residents, which counts as fair, an argument for standardised technique and standardised words.

What the numbers should change in teaching

The conclusion is not to examine less. It is to examine with the numbers in mind, and to teach them.

The stethoscope was invented so that the physician could hear what the body would not say. Two centuries on, the question is no longer whether it hears, but whether two physicians hear the same thing, and the honest answer, sign by sign, is the beginning of a rational examination rather than the end of one.

✦ ✦ ✦

Sources

© 2026 Husain Alkhaldy — Evidence-Based Clinical Examination.