Skip to content

Posts › Development

Development

Five reasons to write a test, with Pest 5

TDD, regression, characterization, acceptance and smoke tests can look identical in a Pest file. What separates them is the reason they were written, and that reason decides how you read them, name them, and whether you can delete them. A field guide with one running example in Pest 5.

14 min read
A field-guide plate of a plant growing from a tests/ pot, its branches labelled TDD, regression, characterization and acceptance

You open a test file in a project you did not write and find it('treats a zero quantity as one'). Is that the specification, or a bug somebody decided to keep? Can you change the behaviour, or will someone in accounting call you on Monday? The answer depends on the reason the test was written, and that reason is rarely in the code.

Most articles about types of tests classify them by scope: unit, feature, browser. In this article I classify them by intent: why the test exists, which is the question above. Scope and technique are separate axes, and I come back to them at the end.

A test can be classified by intent, by scope and by technique: this article is about intent

All the examples grow around one piece of code: a calculator for invoice totals in a small Laravel shop. The same code collects tests for five different reasons, and that is the point. Every snippet runs on Laravel 13 with Pest 5.3.

TDD tests: the test comes first

A TDD test is written before the code it tests, to decide what that code should do. The cycle is the usual one: write a failing test, write the least code that makes it pass, refactor, repeat.

The shop needs to compute invoice totals. Prices are integers in cents, so there is no floating point to worry about. The first test is the simplest thing I can ask:

php
 1use App\Billing\InvoiceCalculator;
 2use App\Billing\Line;
 3
 4it('totals an empty invoice to zero', function () {
 5    expect((new InvoiceCalculator)->total([], euCustomer: false))->toBe(0);
 6});

It fails because InvoiceCalculator does not exist. Creating the class with a total() method that returns 0 makes it pass, and it is the right amount of code for now: the next test will force the rest.

php
 1it('multiplies the unit price by the quantity', function () {
 2    $lines = [new Line(unitPrice: 1000, quantity: 3)];
 3
 4    expect((new InvoiceCalculator)->total($lines, euCustomer: false))->toBe(3000);
 5});
 6
 7it('applies the line discount', function () {
 8    $lines = [new Line(unitPrice: 1000, quantity: 2, discountPercent: 10)];
 9
10    expect((new InvoiceCalculator)->total($lines, euCustomer: false))->toBe(1800);
11});
12
13it('adds 20% VAT for EU customers', function () {
14    $lines = [new Line(unitPrice: 1000)];
15
16    expect((new InvoiceCalculator)->total($lines, euCustomer: true))->toBe(1200);
17});

After a few cycles the calculator looks like this:

php
 1final class InvoiceCalculator
 2{
 3    private const VAT_PERCENT = 20;
 4
 5    /** @param  list<Line>  $lines */
 6    public function total(array $lines, bool $euCustomer): int
 7    {
 8        $subtotal = 0;
 9
10        foreach ($lines as $line) {
11            $gross = $line->unitPrice * $line->quantity;
12            $subtotal += $gross - intdiv($gross * $line->discountPercent, 100);
13        }
14
15        $vat = $euCustomer ? self::VAT_PERCENT : 0;
16
17        return $subtotal + intdiv($subtotal * $vat, 100);
18    }
19}

The VAT rule is simplified on purpose: real VAT rates vary by country and by kind of customer, and one flat rate is enough for the example.

The value of TDD tests is in the design they push you toward. The calculator takes plain value objects and returns an integer, because that is what was easy to test. They also give you a fast loop, and Pest 5 makes it faster: ./vendor/bin/pest --tia records which tests depend on which files, and on the next run executes only the affected ones and replays the rest from cache. After a change to the calculator in the example project:

text
 1Tests:    2 failed, 16 passed (23 assertions, 8 affected, 10 replayed)

Test impact analysis needs PCOV or Xdebug to record the first graph, and it is meant for local development; keep it out of CI.

The limit of TDD tests is that they cover the cases you thought of. All my discounts were round numbers, and intdiv() truncates. I did not think about that, so no test asks about it. If you want a longer worked example of the cycle, I wrote one years ago for a finite state machine in Laravel, with PHPUnit, but the process is the same.

Regression tests: a bug that must not come back

A regression test is written after a bug has been found, to make sure it stays fixed. The order matters: first the test that reproduces the bug, then the fix.

Issue #412 arrives from accounting: an invoice with three items at €3.33 and a 15% discount shows €8.50, while their spreadsheet says €8.49. The discount is 149.85 cents, which intdiv() truncates to 149. Rounded, it would be 150. The test reproduces the report with the same numbers:

php
 1it('rounds the line discount to the nearest cent', function () {
 2    $lines = [new Line(unitPrice: 333, quantity: 3, discountPercent: 15)];
 3
 4    expect((new InvoiceCalculator)->total($lines, euCustomer: false))->toBe(849);
 5})->issue(412);

It fails with exactly the symptom in the report:

text
 1Failed asserting that 850 is identical to 849.

Now I fix the calculator, replacing the intdiv() on the discount with (int) round($gross * $line->discountPercent / 100), and the test passes. A test that never failed proves nothing about the bug: seeing it red first is how you know it reproduces the problem.

The issue() method links the test to the issue. With pest()->project()->github('acme/shop') in tests/Pest.php, the output shows #412 next to the test, and ./vendor/bin/pest --issue=412 runs only the tests linked to it. A year later, whoever reads the test knows why that odd-looking case with €3.33 is there, and where to find the discussion.

Two things make a regression test useful. Name it after the behaviour: it('fixes bug 412') says nothing once the ticket is closed. And put it where the cause is: the bug is in how a discount is rounded, so the test sits on the calculator, below the PDF that displayed the wrong total. Turning bug reports into regression tests is one of the topics I cover at length in Laravel Testing.

Characterization tests: recording what the code does today

A characterization test records what existing code does today, bugs included. You write it when you have to change code you do not understand, typically legacy code without tests, and you need a safety net before touching it.

The shop has an older function, still used by the reports, that computes the totals of imported orders:

php
 1// Inherited from the old shop. Nobody remembers why it does what it does.
 2function legacy_invoice_totals(array $rows, string $zone): array
 3{
 4    $subtotal = 0;
 5    foreach ($rows as $row) {
 6        $price = (float) $row['price'];
 7        if (isset($row['qty']) && $row['qty'] > 0) {
 8            $price = $price * $row['qty'];
 9        }
10        if (! empty($row['promo'])) {
11            $price = $price - ($price * $row['promo'] / 100);
12        }
13        $subtotal += $price;
14    }
15
16    $shipping = $subtotal > 50 ? 0 : 4.9;
17
18    $vat = 0;
19    if ($zone == 'EU') {
20        $vat = round(($subtotal + $shipping) * 0.2, 2);
21    }
22
23    return [
24        'subtotal' => round($subtotal, 2),
25        'shipping' => $shipping,
26        'vat' => $vat,
27        'total' => round($subtotal + $shipping + $vat, 2),
28    ];
29}

I want to replace it with InvoiceCalculator, but I do not know what depends on its quirks. There is no specification to start from, so a characterization test starts from an assertion I know is wrong, and lets the code tell me the answer:

php
 1it('totals two items for an EU customer', function () {
 2    expect(legacy_invoice_totals([['price' => '19.90', 'qty' => 2]], 'EU'))->toBe([]);
 3});
text
 1Failed asserting that two arrays are identical.
 2-Array &0 []
 3+Array &0 [
 4+    'subtotal' => 39.8,
 5+    'shipping' => 4.9,
 6+    'vat' => 8.94,
 7+    'total' => 53.64,
 8+]

I paste the actual value into the expectation and the test passes. Doing this for a handful of inputs exposes the quirks, and each quirk deserves a test whose name describes it:

php
 1it('charges shipping on an order of exactly 50.00', function () {
 2    $totals = legacy_invoice_totals([['price' => '50.00']], 'non-EU');
 3
 4    expect($totals['shipping'])->toBe(4.9);
 5});
 6
 7it('treats a zero quantity as one', function () {
 8    $totals = legacy_invoice_totals([['price' => '12.50', 'qty' => 0]], 'non-EU');
 9
10    expect($totals['subtotal'])->toBe(12.5);
11})->note('Probably a bug. Kept until we know whether anyone relies on it.');

The note() is printed under the test in the output, so the doubt travels with the test.

Hand-written cases cover the quirks I find. To cover what I have not found, I run the function on real orders, anonymised and saved as JSON, and let Pest store the results as snapshots:

php
 1it('matches the recorded totals of real orders', function (string $zone, array $rows) {
 2    expect(legacy_invoice_totals($rows, $zone))->toMatchSnapshot();
 3})->with(fn () => json_decode(file_get_contents(__DIR__.'/../../Fixtures/legacy-orders.json'), true));

Each order in the file has a zone and a rows key, and Pest passes them to the closure as named arguments. The first run writes one snapshot per order under tests/.pest/snapshots and marks the tests as incomplete. From then on, any change to the output fails the test.

The pitfall of characterization tests is fixing things while you write them. An order of exactly €50.00 pays shipping, and a quantity of zero is charged as one. That is how the code behaves, and the test must say so. The decision comes later, quirk by quirk: if the behaviour is wanted, it becomes an acceptance test of the new code; if it is a bug, the characterization test is deleted and a regression test takes its place. Characterization tests are scaffolding: they hold the building up while you change it, and most of them should not survive the refactoring.

Acceptance tests: the requirement in the customer's words

An acceptance test encodes a requirement as an example agreed with whoever asked for the feature. It describes what the system does from the outside, in the customer's terms.

The shop owner's rule is simple: "customers outside the EU pay no VAT". Asked for an example, they give one: a €100 order costs €100 for a customer outside the EU and €120 for a customer in the EU. The test uses those examples and goes through HTTP, the same way a real quote request does:

php
 1describe('a customer outside the EU', function () {
 2    it('pays no VAT on a quote', function () {
 3        $this->postJson('/quotes', [
 4            'eu_customer' => false,
 5            'lines' => [['unit_price' => 10000, 'quantity' => 1]],
 6        ])->assertOk()->assertJson(['total' => 10000]);
 7    });
 8});
 9
10describe('a customer in the EU', function () {
11    it('pays 20% VAT on a quote', function () {
12        $this->postJson('/quotes', [
13            'eu_customer' => true,
14            'lines' => [['unit_price' => 10000, 'quantity' => 1]],
15        ])->assertOk()->assertJson(['total' => 12000]);
16    });
17});

The describe() blocks make the output read like the rule:

text
 1✓ a customer outside the EU → it pays no VAT on a quote
 2✓ a customer in the EU → it pays 20% VAT on a quote

Compare it with the TDD test for VAT. Both check the 20%, but the TDD test talks to the calculator, a class the customer has never heard of, while the acceptance test talks to the endpoint. I can rewrite the calculator, split it, or replace it with a package, and the acceptance test stays the same. That is also why acceptance tests are slower and fewer: one or two per rule, with the edge cases left to the tests closer to the code.

Smoke tests: is it alive?

A smoke test checks that the application starts and that its main pages answer. The name comes from hardware: switch the board on and see whether it smokes.

php
 1it('serves the public pages', function (string $uri) {
 2    $this->get($uri)->assertOk();
 3})->with(['/', '/pricing', '/login'])->group('smoke');

With the smoke group, ./vendor/bin/pest --group=smoke runs them alone, in a fraction of a second, as the first step in CI. A broken configuration, a missing view or a typo in a route fails here before the full suite even starts. Smoke tests are wide and shallow on purpose: a page can answer 200 with the wrong total, and that is a job for the tests above.

Scope and technique

Intent is one axis. The other two are where most of the vocabulary confusion comes from.

Scope is how much of the application a test runs: a unit test calls one class, a feature test goes through the framework, a browser test drives a real browser. Scope is independent of intent. The regression test for #412 is a unit test, but if the bug had been in how the controller reads the discount, the same regression test would have been a feature test.

Technique is how a test checks the result. In this article the characterization tests used snapshots and the rest used plain assertions. Two more techniques are worth knowing, because Pest has them built in.

Architecture tests check the structure of the code:

php
 1arch('billing does not touch HTTP or the database')
 2    ->expect('App\Billing')
 3    ->not->toUse(['Illuminate\Http', 'Illuminate\Support\Facades\DB']);

This one keeps the calculator as pure as the TDD tests made it. If someone later adds a DB::table() call to a class in App\Billing, the test fails with Expecting 'App\Billing' not to use 'Illuminate\Support\Facades\DB'.

Mutation testing tests the tests. Pest changes the code in small ways, a round() becomes a floor(), a > becomes a >=, and runs the tests against each change. If the tests still pass, the change survived, and the behaviour it altered is not really tested. I declared the class under test at the top of the test file with mutates(InvoiceCalculator::class) and ran ./vendor/bin/pest --mutate:

text
 1UNTESTED  app/Billing/InvoiceCalculator.php  > Line 21: RoundToFloor
 2
 3-        return $subtotal + (int) round($subtotal * $vat / 100);
 4+        return $subtotal + (int) floor($subtotal * $vat / 100);

When I fixed #412 I also switched the VAT to round(), but every VAT test uses round numbers, so nothing notices if it truncates. It is the same kind of bug, in a place nobody has reported yet. Mutation testing found it before accounting did. The fix is a test with two amounts whose VAT is not a whole number of cents, one that rounds up and one that rounds down, so that neither floor() nor ceil() survives:

php
 1it('rounds the VAT to the nearest cent', function (int $unitPrice, int $total) {
 2    expect((new InvoiceCalculator)->total([new Line($unitPrice)], euCustomer: true))->toBe($total);
 3})->with([
 4    [1001, 1201],
 5    [1003, 1204],
 6]);

Mutation testing, and how to judge how good a suite really is, is the last topic of Laravel Testing.

Writing the intent down

The five tests in this article look alike: a closure, some setup, an expect(). What separates them is something the code does not say unless you write it down. Pest gives you the places to do it: the test name, issue() for regression tests, note() for characterization tests, describe() for acceptance tests, group() for smoke tests. The next person who opens the file and asks whether a test can be deleted will find the answer there.

If you want to go further, Laravel Testing covers backend testing in Laravel from feature tests to factories, fakes, queued jobs, suite speed and CI, including turning real bugs into regression tests and mutation testing.

Get the next posts by email

An occasional newsletter with my new Posts, with one-click unsubscribe. How I handle your address.