How to Find Out What's Actually Affecting Your Sleep and Mood

7 min read · Updated 2026-07-28

To find what moves your sleep or mood, compare the days a factor happened against the days it didn't, then check whether the gap is bigger than the noise. The catch is that testing many factors at once manufactures false positives, so the comparison needs a multiple-testing correction before you act on anything it says.

You already suspect something is wrecking your sleep. Caffeine after 2pm, the late workout, the second glass of wine, the day you skipped lunch. The suspicion is easy. Confirming it is where people give up, because a line chart of sleep hours next to a line chart of anything else mostly just looks like two wiggly lines.

The method that works is boring and it is not new. Split your days into two piles. Days the thing happened, days it did not. Compare the average outcome across the two piles. That is it. Bearable popularized this for symptom tracking and it is the right shape for the question.

The part almost everyone gets wrong

Here is where it falls apart. Say you track twelve factors and check each one against sleep, mood, and weight. That is thirty-six comparisons. At the usual 5% threshold, you would expect roughly two of them to look significant by pure chance even if nothing in your life affects anything.

So you get an insight card telling you that Tuesdays are ruining your sleep, and it is real in the sense that the math checked out, and it is meaningless in the sense that you rolled the dice thirty-six times and this is what came up. Then you rearrange your week around it.

This is the multiple comparisons problem, and it is the single biggest reason personal-tracking insights feel unreliable. The fix is a correction that adjusts the threshold for how many tests you ran. HabitSync uses Benjamini-Hochberg across every comparison in the batch, not per test. A finding is only labeled strong if it survives that correction and its confidence interval excludes zero.

Why sample size decides everything

A four-day pile against a three-day pile can produce a huge gap and mean nothing. Twenty-five days a side producing a small gap can mean quite a lot. Raw difference is the wrong thing to sort by, but it is what most apps show you, because it makes for a better-looking number.

Ranking on the conservative end of the confidence interval fixes that. If the interval for a factor's effect runs from 0.2 to 1.8 hours, the honest headline is 0.2, not 1.0. And if the interval crosses zero at all, the data cannot even establish the direction, so the finding gets pushed down rather than dressed up.

The practical effect is that a flashy four-day fluke stops outranking a pattern you built over a month.

What you need to log

The comparison needs both piles to exist. That sounds obvious and it is the thing people get wrong, because it is tempting to only log the good days. If you record your workouts but not your rest days, the analysis has nothing to compare against.

Two or three weeks of daily logging is roughly where this starts to be worth reading. Before that the piles are too small and everything comes back weak, which is the correct answer, not a bug.

  • One outcome you genuinely care about. Sleep duration, restedness, or mood are the usual ones.
  • A handful of candidate factors, not everything you can think of. More factors means a harsher correction and more days needed.
  • The boring days. A day with nothing notable is data, not a gap.

Reading the result without fooling yourself

Two habits will keep you honest. First, look at the sample size before the effect size. Second, treat a weak label as a genuine result rather than a broken feature. Weak means the evidence is not there yet, and the useful response is more days, not a different app.

It is also worth muting the pairs you have already chased down. Once you know that late caffeine costs you sleep and you have decided what to do about it, having that card resurface every week just buries the things you have not looked at. Muting a factor and outcome pair, or excluding a metric from analysis entirely, is what keeps the list about new information.

This finds patterns, not causes

A comparison like this cannot tell you which direction the arrow points. Bad sleep might be causing low mood, low mood might be costing you sleep, or a rough stretch at work might be driving both. The statistics do not resolve that and no app can resolve it for you from observational data.

What it does give you is a much shorter list of things worth testing deliberately. Change one, hold the rest, watch what happens over the next couple of weeks. That is a real experiment, and it goes a lot faster when you start from three plausible candidates instead of twelve.

The goal is not certainty. It is spending your attention on the two or three things that have some evidence behind them, instead of the one you happened to notice on a bad Tuesday.

Keep reading