# Goal Architecture for Long-Horizon Self-Study Under Unpredictable Difficulty: An Evidence Review

## TL;DR
- **Keep the permitted-miss mechanic, drop the streak.** The single best-supported change is to replace your Duolingo-style consecutive-day streak with a *cumulative count plus a small budget of pre-committed, mildly costly skip days ("emergency reserves")*. Emergency reserves are the one intervention with direct experimental evidence (one field + four lab studies) that they increase persistence *specifically after a subgoal is missed* — which is exactly your failure mode.
- **Shrink and redefine the unit, but keep a stretch target.** Convert fixed output quotas ("1 chapter," "2 problems") into a **time box with a range goal** (e.g., a 45-minute floor session with a "1–2 problems" range) so a hard unit degrades into "I put in my time" rather than "I failed." Output-unit quotas fail catastrophically under difficulty variance because their cost is unbounded; time boxes and range goals fail gracefully.
- **Adopt a state-based stop rule, not a clock.** The learning-science evidence does not support a magic "20-minute rule" for a hard problem. It supports stopping when your *perceived rate of progress hits zero* (region of proximal learning) or when you hit a genuine *impasse* — then take a hint or a prerequisite detour, because instruction delivered at an impasse is where learning actually happens.

## Key Findings

**On the reader's core premise.** The belief that "daily work is the most effective technique because it keeps you engaged" conflates two distinct objectives that the literature treats separately: **habit automaticity** (which rewards frequency and cue-consistency) and **retention/transfer** (which the spacing literature shows can *favor gaps*). Daily cadence is well-supported for building a *habit* and for engagement/adherence; it is *not* established as superior for *learning outcomes*, and for durable retention, spaced practice beats massed. The reader is partly right (regularity does predict outcomes) but the causal claim is weaker than he thinks and is confounded by motivation. Note the scope limit on the strongest spacing evidence: the landmark synthesis (Cepeda, Pashler, Vul, Wixted & Rohrer, 2006, *Psychological Bulletin* 132(3):354–380) covers **verbal recall over 14,000+ participants**, not complex cognitive skill acquisition, so its transfer to LeetCode/technical reading is itself an inference.

**The mechanism the reader identified is real and correctly diagnosed.** A fixed-unit daily quota implicitly assumes homogeneous per-unit cost. In cognitively demanding domains, per-unit cost is high-variance: most units are cheap, occasional units cost several times the mean, and — critically — the learner cannot forecast which. A quota calibrated to the median unit fails predictably at the tail, and streak/all-or-nothing framing converts each tail event into motivational damage beyond the lost day. Four separate documented mechanisms are doing the damage, and they have *different fixes*: (1) **cognitive load / element interactivity + expertise reversal** (why later chapters overrun); (2) **wheel-spinning / unproductive persistence** (why one hard problem eats the whole budget); (3) **broken-streak demotivation + the abstinence-violation "what-the-hell" effect** (why one miss cascades); (4) **the planning fallacy** (why the day had no buffer for the expensive unit).

**The strongest evidence is cross-domain.** The cleanest experiments on slack, streaks, and rigid-vs-flexible routines come from *exercise, diet, weight-loss, and volunteering* — physical/low-cognitive-load behaviors where per-unit cost is near-constant. This is precisely the transfer risk: a gym visit costs ~the same every day; a chapter does not. Every recommendation resting on that evidence is flagged below.

## Details

### Verdict table

| Design choice | What the evidence supports | Grade | Key citations | Evidence domain | Main caveat for this reader |
|---|---|---|---|---|---|
| **Daily vs. other cadence** | Daily aids *habit/adherence*; for *retention*, spacing (gaps) is superior. Belief in daily-is-best conflates the two. Flexible cadence beat rigid daily-window in the one head-to-head field test. | Moderate | Beshears et al. 2021; Cepeda et al. 2006; Rai et al. 2023 | Exercise; verbal memory; volunteering | No head-to-head daily-vs-weekly RCT on *cognitive skill acquisition* with both adherence and learning measured. |
| **Output-unit vs. time-box vs. process goal** | No direct head-to-head experiment in a variable-difficulty cognitive domain. Inference: output units have unbounded tail cost; time boxes cap it. Learning/process goals beat performance/output goals on *novel complex* tasks. | Weak (unit-type) / Moderate (learning vs performance goal) | Seijts & Latham 2005; Latham, Seijts & Crim 2008 | Management lab tasks | The unit-type comparison is an inference from adjacent literatures, not a measured result. |
| **Rigid all-or-nothing vs. reserve-based slack** | Pre-committed slack *with a small cost* ("emergency reserves") increases both goal preference and persistence, and specifically rescues persistence after a subgoal miss. Flexibility beats rigidity in exercise. | Moderate–Strong | Sharif & Shu 2017; Sharif & Shu 2021; Beshears et al. 2021 | Gym, weight loss, lab goals | All physical/financial/lab; cost must be *small* and reserve framed "emergency" or it licenses slippage. |
| **Single-point vs. range goal** | High–low range goals increase *goal re-engagement* vs. single-number goals, via greater perceived attainability + challenge. | Moderate | Scott & Nowlis 2013 | Weight loss, saving, consumption | Measures re-engagement/likelihood, not learning depth. |
| **Streak-tracked vs. cumulative-count** | Intact streaks boost engagement while intact; a *broken* streak sharply cuts subsequent engagement, worse when self-attributed. "Repair" mechanics attenuate the penalty. | Moderate | Silverman & Barasch 2023; Duolingo (vendor) | Fitness/language apps, games | Outcome is *engagement*, not learning. |
| **Fine vs. flexible subgoals** | Both granular framing *and* flexibility help; the more-flexible granular subgoal produced more *durable* gains. | Moderate–Strong | Rai, Sharif, Chang, Milkman & Duckworth 2023 | Volunteering (N=9,108 field) | Volunteering has near-constant per-hour cost; a chapter does not. |

### Mechanism narrative

**1. Cognitive load, element interactivity, and expertise reversal (the technical-book case).** Later chapters on how LLMs are built have higher *element interactivity* — many concepts that must be held in working memory simultaneously — so intrinsic load rises non-linearly per chapter. This is why "1 chapter/day" works early and breaks late. The **expertise reversal effect** (Kalyuga, Ayres, Chandler & Sweller 2003, *Educational Psychologist* 38:23–31) adds a subtlety directly relevant to a self-learner: instructional support that helps a novice becomes redundant or harmful as expertise grows, and — conversely — a chapter that assumes prerequisite schemas the reader lacks will overload him even if a more expert reader would breeze through. The difficulty spike is *relative to his own prior knowledge*, exactly as he suspected.

**2. Wheel-spinning / unproductive persistence (the LeetCode case).** Beck & Gong (2013, AIED) formalized "wheel-spinning" as failing to reach mastery in a timely manner; the operational thresholds in intelligent tutoring systems are **mastery = 3 consecutive correct, timeliness = within 10 practice opportunities (ASSISTments) or 15 (Cognitive Tutor)** — as operationalized in Kai, Almeda, Baker, Heffernan & Heffernan (2018, *JEDM* 10(1):36–71), wheel-spinning students "practiced the same skill set over 10 times but failed to submit correct answers three times in a row." In Wan & Beck's ASSISTments dataset ("Considering the Influence of Prerequisite Performance on Wheel Spinning," EDM 2015), **20.6% of the 31,301 student-skill pairs in the 2010–2011 training set were classified as wheel-spinning** ("we consider students who fail to achieve mastery within 10 practice opportunities for a skill (including indeterminate cases) as wheel spinning"). The construct is the closest formalization of "spent too long on one hard problem," but the threshold is *opportunities/attempts* on short mastery items, not wall-clock time on one open-ended LeetCode problem — transfer is an inference.

**3. Broken-streak demotivation + abstinence-violation ("what-the-hell") effect.** Silverman & Barasch (2023, *Journal of Consumer Research* 49(6):1095–1117) show across seven studies that broken streaks depress subsequent engagement, and the penalty is *amplified when the person attributes the break to themselves* and *attenuated when consumers can "repair" a broken streak*. This maps onto the relapse-prevention abstinence-violation effect: one lapse triggers all-or-nothing thinking ("I already broke it, so why bother"), and the reader's displacement of *other* commitments after a miss is the classic cascade.

**4. Planning fallacy.** People systematically underestimate task duration (Buehler, Griffin & Ross 1994). The day had no buffer because the plan was built for the median unit. Debiasing that actually works: **unpacking** the task into subcomponents produces longer, less biased estimates, and the benefit grows with task complexity (Kruger & Evans 2004, *JESP* 40:586–598). Counter-finding: task *segmentation* can cause *over*-allocation of time (Forsyth & Burt, in the tradition of the segmentation studies reported in *Memory & Cognition* 36:791–798) — which, for scheduling a day that may hold one expensive unit, is the safe direction to err.

### The "hard unit" decision rule (central to the LeetCode case)

**Verdict: there is no empirically validated fixed time threshold ("20 minutes then look at the solution") in the primary literature.** The widely repeated "20-minute rule" for LeetCode is practitioner lore. What the literature *does* support are **state-based** stop signals:

- **Region of Proximal Learning (Metcalfe 2002, *JEP:General* 131:349–363; Metcalfe & Kornell 2005, *Journal of Memory and Language* 52:463–477).** Skilled learners allocate study time to medium-difficulty (edge-of-competence) items and *stop when their perceived rate of learning approaches zero*. The explicit model: "When the jROLs [judged rate of learning] are zero, people stop studying" — a condition that occurs both when an item is mastered *and* when it is too hard to make progress. **Decision rule: if you cannot detect forward progress after a bounded effort, stop — that is the same signal as having mastered it.** (Paradigm: paired-associate vocab, small-N college samples; conceptual transfer only.)
- **Labor-in-vain effect (Nelson & Leonesio 1988, *JEP:LMC* 14:676–686).** Adults given accuracy emphasis spent substantially more self-paced study time on the hardest items yet showed *no reliable recall gain* — "large increases in self-paced study time can yield little or no increase in the subsequent likelihood of recall." Pouring time into the hardest unit is often literally labor in vain.
- **Impasse-driven learning (VanLehn, Siler, Murray, Yamauchi & Baggett 2003, *Cognition and Instruction* 21:209–249).** From ~125 hours of expert human tutoring dialogue with university physics students: "Successful learning appears to require that the student reach an impasse. When students were not at an impasse, learning was uncommon regardless of the tutorial explanations employed." **Implication: struggle until a genuine impasse — the point where you recognize a specific knowledge gap — then take the hint, because that is when instruction sticks. Taking the hint before the impasse wastes it; grinding long past it is wheel-spinning.** (Association from dialogue analysis, not an experimentally manipulated condition.)
- **The assistance dilemma (Koedinger & Aleven 2007, *Educational Psychology Review* 19:239–264).** The optimal amount of assistance is an inverted-U with no universal setting. In a decade-later retrospective the authors downgraded their own confidence in on-demand hints, finding that feedback on help-seeking "helped students to use on-demand help more deliberately … but not to achieve better learning outcomes" (Aleven, Roll, McLaren & Koedinger 2016, *IJAIED* 26:205–223). Learners are also demonstrably poor judges of when they need help (Aleven & Koedinger 2000).
- **Productive failure (Kapur 2008, *Cognition and Instruction* 26:379–424; Sinha & Kapur 2021, *Review of Educational Research* 91(5):761–798).** Struggling *before* instruction produces better conceptual understanding and transfer than instruction-first: the PS-I meta-analysis reports **Hedge's g = 0.36 [95% CI 0.20, 0.51]**, rising to 0.37–0.58 at high design fidelity, and g ≈ 0.87 after a publication-bias adjustment. But the unguided phase in these studies is a *fixed, teacher-scheduled ~40 minutes* (Kapur 2012) or a multi-lesson unit ending in a consolidation lecture (Kapur 2010) — a design parameter, not an empirically derived optimum, and it *expects* non-solution. Productive failure requires relevant prior knowledge and a *guaranteed* consolidation step afterward; unbounded solo struggle with no consolidation is just failure.

**Synthesized, transferable stop rule for the reader:** struggle on a hard problem until either (a) you reach a genuine impasse you can *name* (a specific missing concept), or (b) your sense of forward progress flatlines — whichever comes first. Then take a targeted hint or do the prerequisite detour, *then* return and re-attempt from scratch to consolidate. Time-box the whole thing to protect the rest of the day, but treat the clock as a *budget cap*, not the learning signal.

### Difficulty calibration (the 85% question)

The **Eighty Five Percent Rule** (Wilson, Shenhav, Straccia & Cohen 2019, *Nature Communications* 10:4646) derives an optimal training error rate of ~15.87% — but *only* for stochastic-gradient-descent learners on **binary classification**. The lead author explicitly declines to generalize it: per University of Arizona News (Nov 5, 2019), "Since Wilson and his collaborators were looking only at simple tasks in which there was a clear correct and incorrect answer, Wilson won't go so far as to say that students should aim for a B average in school." Treat "aim to fail ~15% of the time" as a *loose heuristic*, not a validated target for complex human learning. It converges qualitatively with the region-of-proximal-learning finding that medium-difficulty (not hardest) items are where learning is fastest — the more defensible basis for the reader to select his next item.

### Recovery after a lapse

- **Self-forgiveness (Wohl, Pychyl & Bennett 2010, *Personality and Individual Differences* 48:803–808):** N=119 first-year students; those who forgave themselves for procrastinating before exam 1 procrastinated less before exam 2. **Correlational, single student sample** — suggestive, not causal.
- **Fresh-start effect (Dai, Milkman & Riis 2014, *Management Science* 60:2563–2582):** goal pursuit spikes after temporal landmarks (new week, month, birthday). Archival field studies. Practical use: schedule *re-entry* after a miss to the next natural landmark (Monday) rather than treating the miss as a running deficit.
- **Emergency reserves after failure (Sharif & Shu 2021, *OBHDP* 163:17–29):** the mechanism that directly targets "I missed, so I quit" — reserves preserve a sense of *perceived progress* after a subgoal miss, sustaining commitment (one field study + four lab studies).
- **Implementation intentions / coping planning (Gollwitzer & Sheeran 2006, *Advances in Experimental Social Psychology* 38:69–119):** if-then plans have a medium-to-large effect on goal attainment (**d = 0.65 across 94 studies, n = 8,461**). Second-order ("coping") if-then plans specify what to do *when* the plan is disrupted — directly applicable to "if a problem eats my budget, then I log it and take the hint."

### Is daily best? Adjudicating the premise

- **Correlational study-regularity evidence** (learning analytics / MOOC literature) consistently finds regularity predicts achievement and completion. But this is **confounded**: regular studiers are plausibly more motivated, more self-regulating people to begin with. The MOOC self-regulated-learning literature (e.g., Vilkova 2019, N=2,815; Kizilcec et al. 2017) finds the *forethought* phase — goal-setting, planning, self-efficacy — predicts completion, whereas the performance and self-reflection phases often do not, suggesting the driver is *planning quality*, not raw daily frequency.
- **The one clean causal head-to-head on rigid vs. flexible cadence** (Beshears, Lee, Milkman, Mislavsky & Wisdom 2021, *Management Science* 67(7):4139–4171) found *routine* (fixed daily-window) incentives produced *fewer* gym visits than *flexible* incentives, both during and after the intervention — evidence *against* rigid daily scheduling, in the physical domain.
- **Bottom line:** "daily keeps me engaged" is a defensible *adherence/habit* claim, weakly causal. "Daily is the most effective for *learning*" is **not** established and is contradicted by spacing evidence for retention. Keep near-daily *contact* for habit, but decouple the *learning target* from a daily output quota.

## Evidence tables per sub-question

**SQ1 Cadence.** Beshears et al. 2021 (field RCT, exercise, flexibility > rigidity, post-intervention persistence favored flexible) — *Moderate, physical domain*. Cepeda et al. 2006 (meta-analysis: "839 assessments of distributed practice in 317 experiments located in 184 articles," 14,000+ participants, verbal recall; spacing > massing; "the ISI producing maximal retention increased as retention interval increased," optimal gap ≈ 10–20% of retention interval falling to ~5% at one year) — *Strong for retention, verbal memory only*. **No RCT directly comparing daily vs. every-other-day vs. weekly quotas for cognitive skill acquisition measuring both adherence and learning.** *Absent.*

**SQ2 Unit definition.** No direct experiment comparing output-unit vs. time-box vs. process goal under variable difficulty (*Absent*). Learning-vs-performance-goal literature (Seijts & Latham 2005; Latham, Seijts & Crim 2008; Chen & Latham 2014) shows learning/process goals beat performance/output goals on novel complex tasks — *Moderate, management lab tasks, student/working-adult samples.*

**SQ3 Slack / reserves.** Sharif & Shu 2017 (*JMR* 54:495–509; lab + gym; reserves preferred and increase persistence; "people try to protect their reserve … only using it when they truly must") — *Moderate*. Sharif & Shu 2021 (*OBHDP* 163:17–29; 1 field + 4 lab; reserves rescue persistence after subgoal failure via perceived progress) — *Moderate*. **Moderator: cost must be small and framed "emergency," or slack licenses slippage.** All physical/financial/lab.

**SQ4 Streaks.** Silverman & Barasch 2023 (*JCR* 49:1095–1117; seven studies; streak = 3+ consecutive; intact > broken for engagement; amplified by self-attribution; attenuated by repair) — *Moderate, engagement outcome only*. Duolingo (vendor, non-independent, from Shuttleworth/Lenny's Podcast and company write-ups): equipping two streak-freezes increased DAU by +0.38%; 600+ experiments on streaks over four years; self-chosen goals beat pre-assigned harder goals for retention; loss aversion kicks in ~day 7 — *vendor-reported, engagement not learning*.

**SQ5 Goal type on complex tasks.** Seijts & Latham 2005; Latham, Seijts & Crim 2008 — on novel/complex tasks a *learning* goal (acquire strategy/knowledge) outperforms a *performance* goal (hit a number), because performance goals divert attention to outcome and consume working memory during the declarative stage — *Moderate*. Over-specified quantitative goals' documented harms ("Goals Gone Wild," Ordóñez, Schweitzer, Galinsky & Bazerman 2009) — *cited from secondary sources; original not opened; treat as unverified.*

**SQ6 Range goals / floors.** Scott & Nowlis 2013 (*JCR* 40:444–459; high–low range goals increase re-engagement vs single-number goals, via attainability + challenge → feelings of accomplishment; includes a Weight Watchers field study) — *Moderate, weight/saving domains, re-engagement not learning depth.* Trade-off of shrinking the unit (adherence preserved at cost of learning progress) is **not directly quantified in any study found** — *Absent.*

**SQ7 Granularity vs. flexibility.** Rai, Sharif, Chang, Milkman & Duckworth 2023 (*Journal of Applied Psychology* 108(4):621–634; preregistered field RCT N=9,108 + vignette N=900; granular subgoals — 4h/week or 8h/2 weeks — increased hours volunteered ~7–8% over 12 weeks; the *more flexible* granular subgoal, 8h/2wk, produced more *durable* benefit) — *Moderate/Strong, volunteering (constant-cost behavior).*

**SQ8 Knowing when to stop.** Beck & Gong 2013 + Kai et al. 2018 (wheel-spinning; 3-correct/10-opportunity thresholds; 20.6% of student-skill pairs in one dataset); Metcalfe 2002 & Metcalfe & Kornell 2005 (region of proximal learning; stop when jROL→0); Nelson & Leonesio 1988 (labor-in-vain); VanLehn et al. 2003 (impasse-driven; learning "uncommon" absent an impasse; ~125 hrs tutoring); Koedinger & Aleven 2007 + Aleven et al. 2016 (assistance dilemma; on-demand help "helps but only so much"). **No validated wall-clock threshold; signals are state-based.** *Moderate for the state-based rule; the numeric thresholds are ITS-design parameters.*

**SQ9 Difficulty calibration.** Wilson et al. 2019 (85% rule; error 15.87%; **computational, binary classification, authors decline classroom generalization**) — *Weak for human complex learning*. Converges with region-of-proximal-learning (medium difficulty optimal) — *Moderate.*

**SQ10 High-load content in self-study.** Kalyuga et al. 2003 / Kalyuga 2007 (expertise reversal; prior knowledge × instructional method interaction, technical domains) — *Strong within CLT*. van Merriënboer complex-learning/whole-task sequencing — *Moderate*. Subgoal labeling in programming (Margulieux & Catrambone 2016, *Learning and Instruction* 42:58–71; Margulieux et al. 2020, *IJ STEM Ed* 7 — subgoal-labeled worked examples reduced withdrawal/failure rates in intro programming) — *Moderate, programming education*. Productive failure with guaranteed consolidation (Sinha & Kapur 2021) — *Moderate–Strong*.

**SQ11 Recovery after lapse.** Wohl et al. 2010 (self-forgiveness; correlational, N=119 students); Dai et al. 2014 (fresh-start; archival field); Sharif & Shu 2021 (reserves after failure); Gollwitzer & Sheeran 2006 (implementation intentions d=0.65). *Moderate overall; self-forgiveness evidence is weak/correlational.*

**SQ12 Planning the day.** Buehler, Griffin & Ross 1994 (planning fallacy); Kruger & Evans 2004 (unpacking reduces bias, effect grows with complexity); segmentation studies (over-allocation risk) — *Moderate, mixed samples.*

**SQ13 Premise.** Study-regularity correlational evidence (MOOC/LA) — *Weak causally, confounded by motivation*. Beshears et al. 2021 (causal, against rigid daily) — *Moderate, physical domain*. Cepeda et al. 2006 (against massing for retention) — *Strong, verbal memory.*

## The "hard unit" decision rule — actual thresholds and their transferability

| Source | Threshold/signal actually reported | Design & population | Transfer to solo LeetCode |
|---|---|---|---|
| Beck & Gong 2013; Kai et al. 2018; Wan & Beck 2015 | Wheel-spinning if not 3-consecutive-correct within **10 (ASSISTments) / 15 (Cognitive Tutor)** opportunities; 20.6% of 31,301 student-skill pairs | Log analysis, K-12 math ITS | *Attempts on short mastery items ≠ time on one hard problem.* Inference. |
| Metcalfe & Kornell 2005 | Stop when **judged rate of learning → 0** | Lab, paired-associate vocab, college students (small N) | State signal transfers conceptually; paradigm is memorization. |
| Nelson & Leonesio 1988 | Large extra study time on hardest items → **no reliable recall gain** | Lab, trigrams/general info, adults | Conceptual transfer; memory not reasoning. |
| VanLehn et al. 2003 | Learning "uncommon" **unless at an impasse** | ~125 hrs 1:1 expert tutoring, university physics | Strong conceptual fit; not experimentally manipulated. |
| Kapur 2012 | Unguided phase ≈ **40 min**, then consolidation | Quasi-exp, Singapore 9th-grade math, collaborative | Fixed design parameter, not an optimum; requires guaranteed consolidation. |

## Popular claims without adequate support

- **"21 days to form a habit."** No research basis. The actual empirical study — Lally, van Jaarsveld, Potts & Wardle (2010, *European Journal of Social Psychology* 40(6):998–1009; 96 volunteers over 12 weeks) — found that "the time it took participants to reach 95% of their asymptote of automaticity ranged from 18 to 254 days" (**median ~66 days**), and that *missing a single opportunity did not materially affect habit formation* — a direct empirical rebuttal of all-or-nothing streak logic. The "21 days" figure is self-help folklore.
- **"Don't break the chain / streaks are the best motivator."** Streaks demonstrably raise *engagement* while intact (Silverman & Barasch 2023) but the *broken*-streak penalty is a documented cost, worse under self-blame — a double-edged mechanic, and the outcome measured is engagement, not learning.
- **"20-minute rule for LeetCode."** Practitioner lore. No primary study validates a fixed wall-clock cutoff; the supported signals are state-based (impasse, flatlined progress).
- **"Pomodoro = 25 minutes."** The specific 25-minute interval has no special empirical status; it is a productivity-method convention. What *is* supported is time-boxing to cap the planning-fallacy tail, not the specific number.
- **"Aim to fail 15% of the time" (85% rule) as a study target.** Derived for gradient-descent binary classifiers; the lead author explicitly declines to generalize to classroom learning. Loose heuristic only.
- **Duolingo marketing stats ("streak-7 users 3.6× more engaged," "streak freeze cut churn 21%," "widget +60%").** These circulate on marketing blogs without an independent methods write-up; treat as unverified vendor/marketing claims. The one relatively concrete vendor-reported experimental figure is the **two-streak-freezes → +0.38% DAU** result.

## Contradictions and open questions

- **Spacing vs. habit-frequency.** Spacing says gaps aid *retention*; habit theory says frequency + cue-consistency aids *automaticity*. Reconciliation: these optimize *different outcomes*. For a 3–4 month skill goal you want both — frequent *contact* (habit) but *distributed* re-testing of any specific skill (retention). Not actually in conflict once you separate "show up often" from "re-test the same material immediately."
- **Slack increases persistence vs. slack licenses slippage.** Sharif & Shu show reserves *help* — but only when the cost is small and the reserve is framed as "emergency" (people protect it). Malleable-mental-accounting work shows flexibility can *license* over-consumption. Moderators: **cost of using the reserve** and **whether it's framed as scarce/emergency.**
- **Granularity aids planning vs. flexibility absorbs disruption.** Rai et al. show the *flexible-granular* combination wins for durability — the two are not strictly opposed if you make subgoals granular in *content* but flexible in *timing*.
- **Not studied (honest gaps):** (a) no RCT on daily-vs-other cadence for cognitive skill with both adherence and learning outcomes; (b) no study quantifying the learning cost of shrinking the unit to preserve adherence; (c) essentially all slack/streak/flexibility experiments are physical/financial/short-horizon (2–12 weeks) — none tracks a solo adult learner over 3–4 months on high-load cognitive material; (d) wheel-spinning thresholds are for short ITS items, not open-ended problems.

## Two redesigned protocols

### Protocol A — The technical LLM book (3–4 month horizon)

- **Unit definition:** Replace "1 chapter/day" with a **daily 45-minute time-boxed session** (time box caps planning-fallacy tail — *Moderate, Kruger & Evans*) plus a **content range goal** of "advance 1 section, stretch to 1 chapter" (range goal aids re-engagement — *Moderate, Scott & Nowlis*).
- **Cadence:** Near-daily *contact* for habit, but **re-sequence for spacing**: open each session with a 10-minute recall of a *prior* chapter's concept (spacing aids retention — *Strong for verbal retention, Cepeda; transfer to conceptual material is an extrapolation*).
- **Floor / stretch:** Floor = one 45-min session with any forward motion. Stretch = finish the chapter. Set the floor from your *logged* per-section time distribution (see self-experiment), not optimism.
- **Slack policy:** **2 skip tokens per 2 weeks, mildly costly** (using one requires a 15-minute make-up review next session). Framed as "emergency reserve" (increases post-miss persistence — *Moderate, Sharif & Shu 2017/2021*). *[Extrapolation flag: evidence is gym/weight-loss where a session's cost is constant; a study session's cost is variable.]*
- **Stop rule for a hard chapter:** When a chapter overruns, **do not push through.** Diagnose the impasse (name the missing prerequisite), do a bounded *just-in-time* prerequisite detour, and apply **subgoal labeling** — before reading, write the 3–5 sub-goals the chapter's worked examples accomplish (subgoal labels reduce failure/withdrawal in programming — *Moderate, Margulieux & Catrambone*). If still overloaded, **chunk the chapter into smaller whole-task sections** (whole-task simple-to-complex sequencing — *Moderate, van Merriënboer / CLT*). Expertise-reversal caveat: skip support you no longer need on easy chapters (*Strong*).
- **Tracked/displayed:** A **cumulative "sections completed" count and a calendar heat-map of sessions attended** — *not* a consecutive-day streak (avoids broken-streak penalty — *Moderate, Silverman & Barasch*).
- **Recovery after a miss:** Re-enter on the next **temporal landmark** (next morning / Monday), forgive the lapse explicitly, and use a **coping if-then plan**: "If I miss a day, then I do the floor session next session, no make-up debt beyond one review" (implementation intentions d=0.65 — *Strong; fresh-start Moderate; self-forgiveness Weak*).

### Protocol B — Algorithmic practice ("keep problem-solving sharp")

- **Unit definition:** Replace "2 problems/day" with a **daily 45–60 minute time box** and a **range goal of "1–2 problems"** (range goal — *Moderate*). The output count is a *by-product*, not the target.
- **Difficulty selection:** Pick problems to sit in your **region of proximal learning** — medium relative to *your* current level, not the site's "Medium" label — aiming to solve the majority unaided but be genuinely challenged (*Moderate; loosely consistent with the 85% rule but do not treat 15% as a hard target*).
- **Stop rule (the core fix):** Struggle until **(a) a nameable impasse or (b) flatlined progress**, then take a *targeted* hint (not the full solution), solve it, and **re-attempt from scratch same/next session** to consolidate (impasse-driven learning — *Moderate, VanLehn*; assistance dilemma — *Moderate, Koedinger & Aleven*; labor-in-vain — *Moderate, Nelson & Leonesio*). Hard-cap the single problem at a fixed fraction of the box (e.g., 25–30 min) purely to protect the rest of the day — *this cap is a planning-fallacy safeguard, an extrapolation, not a learning-derived optimum.*
- **Floor / stretch:** Floor = the time box with an honest impasse-and-hint cycle on ≥1 problem. Stretch = 2 problems solved unaided.
- **Slack policy:** Same emergency-reserve structure as Protocol A.
- **Tracked/displayed:** Cumulative problems + **per-problem time logged** (feeds your cost distribution); display a *weekly* count, not a daily streak.
- **Recovery:** Same landmark + coping-plan structure.

## A self-experiment design (n=1)

You are n=1, so use **single-case experimental design (SCED) / N-of-1** methods to get a defensible personal answer.

- **Log from day one — the most important instruction:** per-unit **completion time**, difficulty rating (1–5), whether you hit an impasse, whether you took a hint, whether you hit the floor, and whether you skipped. After ~2–3 weeks estimate your **task-cost distribution** and set the daily floor from a **percentile** (e.g., the 80th-percentile unit time), *not* from optimism. This directly attacks both the planning fallacy and the heavy-tail problem the whole review is about.
- **Design 1 — Alternating treatments** on the LeetCode goal: randomize each week (weekly rather than daily blocks to reduce carry-over) between **(i) fixed output quota** and **(ii) time-box + range goal**. Outcomes: adherence (sessions completed), learning proxy (problems solved unaided on a held-out weekly probe set), displacement (did other commitments suffer?).
- **Design 2 — ABAB** on the streak mechanic: A = consecutive-day streak display; B = cumulative-count + skip-tokens. Watch *re-engagement after a miss* specifically.
- **Design 3 — Multiple-baseline across the two goals:** stagger introduction of the new protocol (book first, LeetCode two weeks later). If behavior changes only when *each* goal's protocol changes, that is strong within-subject causal evidence.
- **Minimum phase lengths:** ≥5 data points per phase; longer given weekend/work-hours noise. Analyze with visual inspection of level/trend plus, optionally, a randomization test.
- **What SCED cannot tell you:** whether results generalize beyond you; it cannot separate maturation (you are getting better at the material anyway) from the intervention unless you use the multiple-baseline; and short phases will not capture the 3–4-month retention outcome you ultimately care about — so keep a delayed retention probe.

## Recommendations (staged)

**Do now (highest evidence-to-effort):**
1. **Kill the consecutive-day streak; switch to a cumulative count + calendar heat-map.** *(Moderate — Silverman & Barasch; removes the self-attributed broken-streak penalty that is your stated cascade trigger.)*
2. **Add a small, mildly costly emergency-reserve of skip days** (2 per fortnight, small make-up cost). *(Moderate–Strong — Sharif & Shu; directly targets "I missed so I quit." Physical-domain caveat.)*
3. **Convert both quotas to a time box + range goal.** *(Moderate — Scott & Nowlis; caps the unbounded tail cost of output units.)*
4. **Adopt the impasse/flatlined-progress stop rule with a time cap.** *(Moderate — VanLehn, Metcalfe, Nelson & Leonesio.)*

**Do next (structural):**
5. **Re-sequence the book:** subgoal-label chapters before reading, chunk overloading chapters, do just-in-time prerequisite detours instead of pushing through. *(Moderate — Margulieux & Catrambone; CLT.)*
6. **Interleave short spaced recall** of prior material into each session. *(Strong for verbal retention — Cepeda; transfer to conceptual material is an extrapolation.)*
7. **Write second-order coping if-then plans** for misses and expensive-unit days. *(Strong — Gollwitzer & Sheeran d=0.65.)*
8. **Re-enter after misses on the next temporal landmark** and forgive the lapse. *(Moderate/Weak — Dai et al.; Wohl et al.)*

**Benchmarks that would change the plan:**
- If your **logged per-unit time distribution is actually low-variance** (tail < 2× median), the premise is wrong for you and a fixed quota is fine — revert.
- If **adherence stays high but weekly unaided-solve rate falls**, you have shrunk the unit too far (adherence bought at the cost of learning) — raise the floor.
- If the **emergency reserve gets used most days**, the cost is too low (it is licensing slippage) — raise the make-up cost or cut tokens.
- If **spaced recall tanks your session enjoyment/adherence**, dial it back — over this horizon, habit maintenance may matter more to you than squeezing out retention gains.

## Caveats
- **Physical-vs-cognitive transfer is the dominant threat.** The strongest experiments (Sharif & Shu, Beshears, Rai, Scott & Nowlis) are gym/diet/volunteering/saving — behaviors with *near-constant per-unit cost*. Your problem is *precisely* that a chapter's cost is not constant. Every reserve/flexibility/range recommendation inherits this caveat.
- **Engagement ≠ learning.** Streak, gamification, and Duolingo evidence measures app-opens/adherence, not skill gain. A vendor optimizing daily active users is not optimizing your understanding of LLM internals.
- **Horizon mismatch.** Most slack/streak/self-forgiveness studies run 2–12 weeks; your horizon is 3–4 months. Two-week adherence effects may not survive it. No cited study tracks a solo adult over your horizon on high-load cognitive material.
- **Publication bias and effect inflation** are known concerns in goal-setting, gamification, grit, and self-control literatures; productive-failure effect sizes shift substantially with correction method (g=0.36 raw vs 0.87 bias-adjusted) — read the bracketed intervals, not the point estimates.
- **Student samples.** Much of the learning-science evidence uses children/undergraduates in classrooms; you are a working adult with constrained evening hours and an exogenous shock (long office hours) that none of these studies model.
- **Unverified items not relied upon for specific numbers:** Ordóñez et al. "Goals Gone Wild"; gamification meta-analyses (Sailer & Homner; Hanus & Fox); Winters & Latham. Where mentioned, they are flagged as to-be-verified and no effect sizes are drawn from them.
