Methodology

An assessment that cannot explain how it was built should not be trusted with anything that matters.

This page explains where the variables come from, how self report is handled, what happens when the data contradicts itself, what the synthesis adds, what the method cannot currently do, and how it will be validated.

We publish the limitations alongside the claims. Most assessment tools do not. That difference is the point.

01

Where the variables come from

The Human Operating System is a structured map of how a person functions across four systems, physical, emotional, cognitive, social, separated into layers: what you started with, what you learned, what drives you, and what you can train.

The variables in that map come from three sources.

Established research traditions. Each variable is anchored to a body of work where the underlying construct already has a track record: temperament and arousal research, sensory processing sensitivity, affect intensity and emotional reactivity, working memory and executive function, attachment theory, schema and cognitive behavioural models of belief formation, self determination theory and motivational research, and behavioural expression models. We are not inventing new human dimensions. We are organising known ones into a structure that makes their interaction visible.

A structural rule about explanation. Every variable has to sit in the layer that actually explains it. This sounds abstract; in practice it is the most consequential design decision in the system. Sensitivity to rejection is constitutional, it belongs at baseline. The strategy someone builds to avoid rejection is learned, it belongs one layer up. Confusing the two produces advice that tries to train away something untrainable, or treats a changeable habit as fixed identity. Several variables have been moved between layers during development for exactly this reason: stress response bias, trust baseline, and threat versus opportunity orientation all began as baseline variables and were relocated once it became clear they describe learned interpretation rather than biological default.

Inclusion criteria. A candidate becomes a variable only if it passes four tests.

  • Distinguishable. It does not collapse into its neighbours. Rejection sensitivity and criticism sensitivity feel similar and are measured separately because they produce different behaviour and respond to different interventions.
  • Mechanistic. There is a plausible account of what produces it, not just a label for a cluster of behaviour.
  • Consequential. Knowing the score changes what we would recommend. If a variable cannot alter a decision, it is description, not measurement.
  • Layer clean. It belongs to exactly one explanatory layer.

What this does not yet mean. The variable set is theory derived and practice refined. It has not been empirically derived through factor analysis on a large sample. That work is scheduled, and it may show that some variables should merge and others should split. Section 05 covers this honestly.

02

How self report is calibrated

Everything in the assessment comes from what the person tells us. The calibration is in the order it is gathered, the form the questions take, and what each answer is checked against.

Context before variables. The process opens with an intake and a full context layer, life history, current situation, the specific areas of the person's life, and the patterns that have run through each of them over time. None of this is scored. It exists so that no variable is ever read in isolation. By the time any individual trait is discussed, there is already a documented account of how this person has actually lived, and every answer that follows sits against it.

Description rather than rating. For each variable, the person is not asked to place themselves on an abstract scale. They are given detailed descriptions of what each expression of that variable actually looks like in lived experience, how it shows up, what it feels like from the inside, what it produces in daily life, and they identify which one matches them.

This matters more than it sounds. Asking someone to rate their sensory sensitivity from one to five is asking them to compare themselves to a reference point they do not have, and the number that comes back is mostly noise. Recognition works differently and better. People are unreliable at quantifying themselves and comparatively good at reading a specific description and knowing whether it is theirs. The format also surfaces distinctions people would never have drawn unaided, the difference between being drained by social contact and being uninterested in it is invisible on a scale and obvious across two well written descriptions.

Every answer has a reference frame. Because the context layer comes first, a self identification is never the only thing on the table. Where someone selects a description that does not fit the life already documented, the gap is visible immediately and gets examined rather than recorded.

What this depends on. Two things, stated plainly.

The first is awareness. Some patterns are structurally difficult to see from the inside, particularly protective ones, which work precisely by not being noticed.

The second is honesty, and this format raises the stakes on it. A detailed description makes clear what choosing it implies. Where a scale item can often be answered without quite knowing what is being revealed, a described way of being can be recognised and declined. Someone who wants to present a preferred version of themselves has an easier time doing it here than on a questionnaire. Nothing in the method removes that. What it does is make the discrepancy findable, because a description chosen for how it sounds rather than how it fits tends not to survive contact with the person's actual history.

Uncertainty is worked, not labelled. The assessment does not attach a confidence rating to an unclear result and move on. If something does not resolve, the description does not quite fit, two seem equally true, the answer sits oddly against the context layer, that is unfinished work, not a finding with an error bar on it. It goes back: the variable is approached from another direction, the concrete instance is asked for, the answer is checked against what the person has already described about their life. That continues until the reading is clear. What appears in the profile is what came through that process, not an average of everything said along the way.

State separated from trait. A person in the middle of a crisis will read low on clarity, low on energy stability, and high on threat orientation, and none of it describes their baseline. Current load is established first, and baseline readings are interpreted against it. Sensitivity variables carry a stress modifier: what this looks like at rest, and what it looks like under pressure. Where those diverge sharply, the divergence itself becomes part of the profile.

03

How contradictions are interpreted

Contradictions are not measurement failure. They are usually the most informative part of the data, and the method is designed to preserve them rather than average them away.

There are two kinds.

Contradiction inside the data. What the person identifies with on one variable does not fit what they identified with on another, or does not fit the life history already documented. These are never averaged into a middle reading. They are checked, directly, in the conversation, until it is clear what is producing the mismatch. The usual explanations:

  • Ideal self versus actual behaviour. The description chosen is the one the person is trying to live up to. Common, and not always dishonesty.
  • Genuine context dependence. The pattern is real in one domain and absent in another. Someone can be highly conflict avoidant at home and direct at work. This is a finding, not an inconsistency.
  • State overriding trait. Current conditions are producing a reading that will not replicate in six months.
  • An awareness gap. The behaviour is visible from outside and not from inside.

Which explanation applies is tested in conversation against real examples, not decided by the scoring rule.

Contradiction between the profile and the person. The harder case: the assessment indicates X, and the person says they are clearly Y. There are four possibilities, and none of them is assumed by default.

  • The measurement is right and the person cannot see it. Most likely for learned protective patterns, which are by definition outside awareness, and for baseline traits currently masked by a long standing compensation strategy.
  • The person is right and the measurement is wrong. Happens when a description is poorly worded, culturally loaded, or ambiguous. This is not settled by assumption in either direction, it is checked, from more than one angle, until it is clear which reading holds. Every disagreement is also logged against the specific description that produced it, and descriptions that generate disagreement across many people are rewritten or removed.
  • Both are partly right. The person is describing a real version of themselves, their stressed self, their recovered self, their self in one specific relationship, that differs from the one being measured. The task is identifying which state each side is describing.
  • The variable is poorly defined. If many people disagree on the same variable, the variable is probably measuring two things at once, and it goes back for structural review.

Two working principles hold across all four. First, the person's perception is never discarded, it is data about self concept regardless of whether it matches the behavioural reading, and a persistent gap between how someone sees themselves and how they operate is itself one of the more useful findings the assessment can produce. Second, a disagreement is not left standing as a disagreement. It is worked, asked again differently, tested against concrete instances, held up against the context layer, until there is full clarity about which reading is accurate and why. A contradiction that has been noted but not resolved is not a result.

04

What the synthesis contributes

Individual variables are weakly informative. Almost everything useful lives in the interaction between them, and the synthesis layer is where the assessment stops describing and starts explaining.

Combinations, not scores. High cognitive load sensitivity on its own is unremarkable, a great many people have it and function well. High cognitive load sensitivity combined with low energy stability, high evaluation sensitivity, and a learned pattern of over preparation is a specific and predictable architecture: the person compensates for load sensitivity by over preparing, over preparation consumes the energy their unstable baseline cannot spare, and evaluation sensitivity prevents them from delivering anything unfinished, so the cycle cannot be exited by effort. That is a mechanism with an intervention point. Neither variable produces it alone.

Causal chains. The layered structure allows a behaviour to be traced downward to what generates it, and, more importantly, to identify which layer change is actually available in. You do not train away a baseline; you design an environment around it. You can update a learned pattern, but only by addressing what it protects. Execution capacity is genuinely trainable. Most self improvement failure comes from applying the wrong one of these three to the wrong layer, and the chain is what prevents that.

Cross system interaction. The strongest findings usually cross boundaries. High social motivation combined with high social load sensitivity produces an approach withdrawal conflict that neither variable predicts by itself: the person wants contact and is depleted by it, and reads their own exhaustion as a lack of interest. High pattern recognition combined with high unclarity sensitivity produces someone who sees implications early and cannot tolerate holding them unresolved. These configurations are invisible in any single system reading.

Leverage points. Because the variables are not independent, most of a profile is downstream of a small number of nodes. The synthesis identifies where a change propagates furthest, usually two or three places, not twenty.

Translation into daily life. The final output is not a profile but a Life Operating System: what each finding means across the specific areas of a person's life, what the aligned and misaligned expression of each pattern looks like, and what to change in routines, environment, and decision rules. A profile that does not reach this stage is interesting and useless.

A note on structure. The assessment is administered in layer order, because that sequence prevents later answers from contaminating earlier ones. The report is delivered by system, physical, emotional, cognitive, social, because that is how people actually think about themselves. The synthesis then reconnects them.

05

Current methodological limitations

This section is deliberately specific. Anything not listed here as established should be assumed to be pending.

There is no normative sample. Scores are interpreted relative to the rest of an individual's own profile, never as a percentile or a comparison to a population. We do not say higher than average. We do not have an average.

This is not a clinical instrument. It does not diagnose, screen for, or rule out any condition. Where responses suggest something that warrants clinical attention, the response is a referral, not an interpretation.

Psychometric properties are not yet established. Internal consistency, factor structure, test retest stability, and differential item functioning across groups have not been tested at the sample sizes required to draw conclusions. The variable set is theory derived. Empirical analysis may show that some variables collapse into each other and others need to split.

Current evidence is predominantly self report. Observer input is available and encouraged but not yet systematically collected, which means blind spot detection currently relies more on cross layer contradiction than on outside confirmation.

Interpretation involves practitioner judgment. Inter rater reliability, whether two trained assessors reading the same responses would reach the same conclusions, is currently unknown, and this is a real limitation for anything delivered in conversation rather than automatically.

Fairness has not been formally tested. The framework has been developed and used with a limited demographic and linguistic range. Whether items function equivalently across cultures, languages, neurotypes, education levels, and socioeconomic contexts is an open question, not a settled one. Some items are known to be culturally loaded and are flagged as such.

Breadth constrains depth. With a large baseline variable set, not every variable currently rests on the full five evidence layers. Those that do not are reported at correspondingly lower confidence, but this means some parts of a profile are considerably better supported than others, and the report says which.

It does not reduce to one number. The framework is a multidimensional composite, and a single overall score would be statistically meaningless. This is a design decision, not an omission, and it means the system cannot be validated as one instrument, each construct requires its own evidence.

Predictive validity is untested at scale. We can show that people recognise themselves in the results. We cannot yet show, with the sample size required, that the results predict outcomes over time. That study is described below.

06

Validation and development roadmap

Validation happens at four levels. We are transparent about which are complete, which are in progress, and which are years out.

Level 1, does it feel accurate. In progress. Recognition rate and accuracy ratings collected immediately after every assessment, plus specific agreement and disagreement per variable. Target: 80 percent or more of participants rating the result as highly accurate. Face validity is the minimum threshold for credibility, not evidence of validity, a well written horoscope also passes it. It is necessary and not sufficient.

Level 2, does it predict anything. In progress. Structured follow up at 4 to 6 weeks and again at 6 to 12 months. Which findings proved accurate in the person's life. Which did not. What they initially rejected and later recognised. This is the level that separates a description from a measurement, and it can be reached without academic infrastructure, it requires only discipline and time.

Level 3, do the variables correlate with established measures. Planned. Whether our sensitivity variables track validated sensitivity scales, our affect variables track established affect intensity measures, our relational variables track validated attachment instruments. This requires a sample of 100 or more and a research psychologist partner, and it is the point at which the framework earns the right to describe itself as measuring what it claims to measure.

Level 4, are variables that should be different actually different. Planned. If rejection sensitivity and criticism sensitivity are genuinely distinct, scores on them should vary independently rather than move together. Wherever they do not, the variables merge. Same infrastructure as Level 3, and it is the test most likely to shrink the framework.

Running in parallel: expert review of items by qualified reviewers; cognitive interviews with test users to check that questions are understood as intended; item level tracking of confusion, skipping, and disagreement; and revision or removal of items that fail.

What we commit to. The framework is versioned and revisions are published, including what changed and why. We will not claim norms without a norm sample, will not claim clinical validity without clinical validation, and will not market a single alignment score. If validation shows that part of the structure is wrong, that finding gets published too.

No label without evidence. No score without uncertainty. No insight without context. No depth without safety.

This assessment is a developmental reflection tool, not a clinical diagnostic instrument. Results are hypotheses based on your responses, not fixed truths. The most useful findings are those that match repeated patterns in your actual life and support constructive action.