puzzle #22
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Optimal house-selling

You will learn: Compute finite-horizon acceptance thresholds from continuation values.

Start with: Expectation and variance

An offer arrives. Accept it now, or pay to wait for another? A high offer can still be worth rejecting when many opportunities remain, while an ordinary offer can be worth taking near the deadline. The comparison is between the current offer and the expected value of continuing from the current decision.

We use a finite token model: independent offers of 100, 150, or 200 tokens from a known distribution, at most n offers, no recall of rejected offers, and a compulsory acceptance of the last offer. The first observation is free. Every rejection costs c tokens. The objective is expected accepted tokens minus rejection costs.

Compare current and future value

Offers are 100, 150, or 200 tokens with probabilities 0.333333, 0.333333, 0.333333. Offers are independent. The first observation is free; each rejection costs 10 tokens. No recall; the last offer must be accepted. Equality with continuation value means accept.

Expected optimal value with more offers available05010015020011.522.533.544.55offers remainingexpected net tokens

Optimal expected net: 175.061728 tokens. Accepting the first offer gives 150.000000 expected tokens. These are expectations before observing any offer.

Offers remainingExpected optimal valueContinue valueOffer 100Offer 150Offer 200
1150.000000Last offer: must acceptAcceptAcceptAccept
2163.333333140.000000RejectAcceptAccept
3168.888889153.333333RejectRejectAccept
4172.592593158.888889RejectRejectAccept
5175.061728162.592593RejectRejectAccept

Follow the policy without seeing future offers

DayOfferActionNet if accepted now
1100Reject and pay cost100
Exact stopping-day probabilities
DayP(stop on day)
10.33333333
20.22222222
30.14814815
40.19753086
50.098765432
A finite toy model with known offer probabilities, linear token utility, and a compulsory terminal sale. Continuation values are computed backward, without looking at the sampled future. Changing the seed changes the example, not the policy.

The probability control sets the chance of 200 tokens. The remaining probability is split equally between 100 and 150. The policy table uses the same distribution on every day. The seed chooses one example sequence; the thresholds are calculated without inspecting it.

Reveal the sequence one offer at a time. The table reports what would remain after costs if that offer were accepted now. Previously paid costs matter for the eventual total, but are sunk at the current decision. Paying them does not make a weak future offer intrinsically better.

Work backward from the last offer

Let Vₜ be the optimal expected value before seeing an offer when t offers remain, excluding costs already paid. With one offer left, acceptance is mandatory, so V₁ = E[X]. With at least two offers left, observing x gives two choices: accept x, or reject, pay c, and obtain the expected continuation value Vₜ₋₁. Therefore

Vₜ = E[max(X, Vₜ₋₁ − c)].

The acceptance threshold is Vₜ₋₁ − c. Accept at or above it; reject below it. Equality is resolved by accepting, which preserves the same expected value and avoids an unnecessary extra observation. On the terminal day the rule is compulsory acceptance, not comparison with a fictional future value.

For the default equal probabilities, V₁ = 150. With two offers left and c = 10, continuation is worth 140. An offer of 100 should be rejected; 150 and 200 should be accepted. Hence V₂ = (140 + 150 + 200)/3 = 163⅓. With three offers left, continuation is worth 153⅓, so both 100 and 150 should now be rejected. The changing threshold follows directly from the remaining opportunities.

Expected payoff is not a promised sale

The blue curve gives expected value before observing the next offer. One sampled path may finish above or below it. To get the exact stopping-day distribution, begin with probability one of reaching the first day. At each day, multiply the surviving probability by the chance that an offer meets that day’s acceptance rule. Carry the remaining mass into the next day. The final day absorbs all remaining mass.

This forward probability calculation and the backward value calculation describe the same policy from different directions. They should agree on its expected net payoff. Neither calculation is allowed to optimize after seeing which future offers actually appeared; that would assign the seller information the model never grants.

Change one assumption at a time

With just one offer there is no decision. When rejection costs are high enough, immediate acceptance is optimal. With zero rejection cost, waiting for a maximum offer is attractive while time remains; nevertheless, a finite horizon still forces acceptance at the end. When the high-offer probability is one, every offer is 200 and accepting immediately is optimal, including at a zero cost tie.

Make a prediction

Would these same thresholds apply if a rejected offer could be recalled later?

Explore the answer

No. The best offer seen so far would then be part of the decision state. Continuing could preserve that option, so the recurrence and thresholds would change. Likewise, learning an unknown offer distribution or maximizing a nonlinear utility requires a different state or objective.

Ferguson’s Optimal Stopping and Applications, chapter four, develops selling and observation-cost models. The experiment here explicitly uses a finite horizon, a free first observation, and costs per rejection. These are defined teaching assumptions, not estimates of a housing market. Continue to the secretary problem to contrast known token values with choosing the best relative rank.

Reset all settings