• ➤

    Picking the next message for one user looks like a ranking problem, but it isn’t quite one. The candidate set changes daily. The best answer differs from user to user. You only ever see the reward for the action you actually took, and it’s sparse even then. And every choice has to respect hard business constraints. That makes it a constrained contextual decision problem, and the stack is built around that.

    CleverAI™ has two coupled halves. The decisioning engine picks the action for each user. The agentic layer, AI Studio, is where agents generate the candidates, run the decision through campaigns and journeys, and keep the whole thing aimed at the business goal.

    The model has to do two jobs. Offline, it learns long-horizon behavior: purchase cadence, lifecycle stage, years of campaign response. Online, a second learner handles what the offline model couldn’t have seen, like drift, a brand-new session, a new category, or a sudden drop in engagement. It updates on every outcome, but it can’t throw away the prior, and it can’t overreact to one event.

    The system side has three hard requirements. Context changes with every event, across hundreds of millions of users, and it has to be readable faster than the request that needs it. Every request evaluates a large candidate space against hard constraints. And the reward loop has to be fast enough that the policy reflects the last interaction rather than the last retrain.

    The agentic layer adds a fourth requirement. Every agent action must be governed, and every agent must draw its decisions from the same policy. Otherwise intelligence fragments across agents.

    The Decision Problem

    At each decision the engine observes a user context and a set of eligible actions . Each action is a campaign together with its delivery choices: channel, timing, offer, or no action at all. The engine picks and later observes a reward (positive if the user moved toward the goal, negative if they moved away). We want a policy that maximizes expected reward while never placing probability on an infeasible action:

    Equivalently, the policy should minimize regret against the best feasible action for each user:

    Three things make this harder than the textbook version. The action set changes every day as campaigns launch and expire, so nothing can be learned once and kept. The best action depends on , so there is no single winner. And feedback is partial: you see the reward for the action you took and never for the ones you didn’t.

    Most of the tools teams reach for first miss at least two of these:

    • Cohort rules give everyone in a segment the same action, ignoring how varies inside it.
    • Sequential A/B tests split traffic uniformly for a fixed period. With variants, a share of traffic goes to losing arms for the whole test, so regret grows linearly in . The output is also one global winner, so even after the test you are still linearly far from the per-user optimum. Adaptive policies achieve regret that grows sublinearly, on the order of up to dimension and log factors.
    • A propensity model on its own estimates : who is likely to act. It says nothing about : what to send them.

    Key Components of the CleverAI™ Live Architecture

    Each component owns one term of the problem above.

    • TesseractDB™ -> context : One real-time feature store holds live event state and up to 10 years of history per user. Both learners read from it, so there is no sync boundary between recent and historical context.
    • Creators and Experts –> actions and their descriptions : Creators supply content variants with embeddings; Experts supply specialized scores for recommendations, predictions, timing, channel, and sequence.
    • Guardrails -> feasible set : These are hard constraints applied inside the decision step, not a post-filter. What comes out is the best valid action, not the best action that happened to survive filtering.
    • The Tesseract Decisioning Engine™ -> policy : an offline-trained prior that bootstraps an online learner updating per interaction.
    • Experience Builder -> executes and returns the reward .
    • The Clean Room -> extension of : Embeddings are computed inside the customer’s environment, so regulated attributes inform ranking without leaving it.
    • AI Studio -> the agentic layer above all of this: hosts the agents and agent harnesses. Agents read and write TesseractDB™, direct Creators and Experts to produce candidates, call the Tesseract Decisioning Engine™ for the selection, and drive Experience Builder to execute. Agent actions are gated by human-in-the-loop (HITL) approvals and RBAC.

    Inputs: Goal, Strategy, Guardrails

    The control plane turns marketer intent into the two objects the engine optimizes: a reward and a feasible set.

    Goal → reward. A marketer names positive signals (checkout, add-to-cart) and negative signals (uninstall, subscription cancellation). The reward for a decision is a function of the goal events that follow it within the goal’s horizon :

    where is the time of the decision. rewards positive events and penalizes negative ones.

    Guardrails define the feasible set. Eligibility, audience boundaries, and frequency and touch-point caps combine into the set of actions the engine may choose from for this user, now:

    Here is the set of caps that apply to action (per channel, global, by team or label) and is the user’s current count against cap in its window.

    These inputs fan out to two consumers. TesseractDB™ receives the audience boundary and goal events, so segments and derived features are computed against them. AI Studio receives the full goal, strategy, and guardrail set as the objective its agents plan against.

    AI Studio

    AI Studio is the agentic automation layer. Its one architectural rule: agents plan and act, but they never make the per-user selection themselves. Every agent that needs a decision calls the same policy , so there is one learned model of the user, not one per agent prompt.

    • Agents: Goal-driven processes that own an outcome. Lifecycle Agents optimize content and sequence toward a milestone like first purchase or KYC completion. One-time and drop-off agents cover calendar sends and funnel abandonment. Teams with a use case we haven’t built can write their own on Foundry and run it on the same infrastructure, with the same policy behind it.
    • MCP Server & Skills: The tool surface agents call: read tools (analytics, segments, campaign history), write tools (create or modify campaigns, journeys, segments), and evaluate tools (performance reads, estimates). The same surface is exposed externally, so a customer’s own assistant can drive the platform through the same contract.
    • Controls: Every agent write goes through a HITL approval gate, and RBAC limits what each agent can read or change.

    Data Layer: TesseractDB™ as the Feature Store

    TesseractDB™ (12+ patents) exists to compute correctly and fast. For user at decision time , the context is a function of that user’s full history in one store:

    The history includes live behavior as it arrives, up to 10 years of event-level history (not sampled or rolled up), current profile attributes, every prior delivery and its outcome, and computed features.

    The defining choice is that lives in one store. Traditional architectures split history (warehouse or CDP) from live state (operational store), which forces a choice between a network join at request time and stale context. Because both learners read from TesseractDB™, the outcome of decision is written back and is part of for decision .

    Clean Room: Extending the Feature Space Without Moving Data

    Internal risk scores, transaction-level attributes, and records under residency constraints are often the strongest signals a customer has, and they cannot leave the customer’s environment. Let be those private attributes. The Clean Room computes an embedding inside the customer’s environment and returns only the embedding:

    is trained to be predictive for the goal while being non-recoverable from .

    Candidate Generation: Creators and Experts

    The engine never scores a campaign ID. It scores a description of the action, which is what lets learning transfer across campaigns. Two upstream layers, coordinated through AI Studio, build that description.

    Creators generate content variants – copy, images, templates, offers, experience variants – within brand rules. Rather than hand-tagging campaigns with a fixed schema, CleverAI™ embeds campaign content into a learned semantic space that captures tone, topic, offer type, urgency, and visual character, then joins that with structured metadata and the delivery choice:

    where is the semantic intent embedding and is structured metadata (audience definition, goal type, time window).

    Experts are specialized models for recommendations, predictions (propensity to convert, churn), best send time, best channel, and sequence intelligence. They are not separate decision engines. Each contributes a score or that the engine consumes as a feature on the user side or the action side.

    Decision Layer: Tesseract Decisioning Engine™

    The engine takes five inputs per request: context (from TesseractDB™ and the Clean Room), choices (from Creators), expert scores, the goal, and the feasible set and emits one output: a best valid action, which may be a message, offer, channel, timing, sequence step, product experience, or no action.

    Two Learners: System 1 and System 2

    Kahneman described human judgment as two systems. System 1 is fast, automatic and built from long experience. System 2 is effortful and engaged when intuition isn’t enough. The engine is built the same way. The offline learner is its System 1: an instant read of who the user is, learned from years of behaviour and retrained only periodically. The online learner is its System 2: it weighs the candidates, spends effort where the answer is unclear, and changes its mind after every outcome. Intuition is fast to use and slow to change; deliberation is effortful and quick to adapt. The expected reward of action for user is written as

    where is the user’s propensity to reach the goal and is a learned representation of the user, both from the offline learner, and is the online learner with parameters .

    and capture who the user is. That changes over weeks and needs long history and a heavy model. captures which option works for them. That changes every time a campaign launches or is edited, and needs a light model updated per interaction. System 1 supplies the intuition; System 2 makes the call.

    System 1: offline learnerSystem 2: online learner
    Estimates and for every
    Depends onUser only, campaign-agnosticUser × action
    Learns fromLong behavioral historyThe stream of
    Update cadencePeriodic batch retrainingIncremental, per outcome

    The analogy is structural, not literal. Both systems run on every request in milliseconds; “slow” refers to how System 1 is retrained, not to how long it takes to answer. And unlike Kahneman’s System 2, ours is not lazy: it is consulted on every decision.

    System 1, the Offline Learner: Understanding the User

    The offline learner is an extension of CleverTap’s production prediction system. It is a non-linear model over profile attributes, behavioral actions, and expert scores, a user-side space that runs into the thousands of dimensions after categorical expansion:

    It is trained offline on long-range history and provides the prior that bootstraps the online policy, so day-one decisions start from learned patterns rather than uniform exploration. A learned feature-selection step prunes the expanded space before training, and its top-ranked attributes are reused as part of the online learner’s user context.

    Framing the label so it doesn’t leak. For a reference time , features summarize a lookback window and the label looks only at the period after it:

    The horizon comes from the business goal, and the lookback is set per goal. Examples are drawn from several reference times rather than one, so the model learns patterns that hold across seasons and campaign cycles rather than the quirks of a single week.

    Separating the user from the old policy. Historical logs contain delivery decisions (when and how each user was contacted), and those were chosen by a past policy . A model of trained on these logs confounds who the user is with how the old policy treated them. So is estimated on user features only. Timing and channel go into , where the online learner treats them as choices to optimize rather than facts to explain.

    The same logic applies to sample selection: if a user’s inclusion in training depends on their behavior, the model learns and is then asked to score everyone. Training samples are drawn to avoid that gap.

    System 2, the Online Learner: A Contextual Bandit Over Described Actions

    The online learner is a contextual bandit. It learns the association between the user-side vector and the action-side vector and produces a score for whether a specific user will respond to a specific candidate.

    System 1 informs System 2. A linear learner over raw user features can only represent effects that add up feature by feature. The offline model has already learned the non-linear structure of user behavior, so the online learner consumes that structure through instead of relearning it. Anything linear in can express patterns the offline model found, while the online model stays small enough to update on every outcome.

    Personalization lives in the interaction terms. The user side stacks , , recent engagement, and selected attributes. The score has main effects plus bilinear user × action terms:

    The main effects alone learn “this campaign is good” and “this user is engaged”, which is not personalization. The bilinear terms learn how a kind of user responds to a kind of creative, channel, or timing. is a deliberately chosen set of user-group × action-group pairs; choosing it is a first-class modeling decision that bounds model size and keeps learning on interactions that matter.

    Learning from partial feedback. Each decision reveals the reward of one action only. The learner updates from the observed decision and its outcome, with a loss that depends only on that record, and generalizes to untried actions through the shared feature space. It does so after every outcome, whether the user converted, ignored the message, or triggered a negative goal event such as an uninstall:

    Exploration that follows the scores. The engine does not always serve the top-scored action. The policy is a probability distribution over that is monotone in the score: clear leaders get most of the traffic, close contenders share the remainder, and clearly weak options get almost none. Starting warm, not cold. A bandit that starts at θ = 0 spends its early traffic relearning what the business already knows. Before serving, θ is warm-started from historical decisions and outcomes, so online exploration refines a sensible prior instead of building one.

    AI Experimentation and Cold Start

    Conventional A/B testing is sequential and memoryless: one hypothesis, one population-level winner, linear regret for as long as the test runs, and nothing carried into next quarter’s version of the same campaign. The engine instead runs the matrix continuously and in parallel, and because candidates are described by rather than enumerated by ID, a new campaign does not start from zero.

    Consider only the user × intent term of the score. For two campaigns and shown to the same user, Cauchy–Schwarz gives

    So for the same user, campaigns with similar intent are guaranteed to get close intent-driven scores, before the new one has a single impression, much as System 1 judges new things by their resemblance to familiar ones. A new campaign inherits a prior from its semantic neighbors, and online updates refine it from there.

    In practice, recurring themes like cart reminders, festive discounts are exploited from the first request. Exploration is spent only where the engine is genuinely uncertain like a new theme, an untested channel-context pair, a segment that has never seen this offer type.

    Constraint Handling

    Guardrails enter the decision as the support of the policy, not as a filter on its output. Scores are computed and normalized over the feasible set only:

    This is why the output is a valid action by construction. It also matters for learning. If a downstream filter discarded the sampled action and sent another instead, the policy would be updated on outcomes of actions it never chose. Constraining before scoring keeps what the engine chose and what the user received the same thing.

    Execution Layer: Experience Builder

    The engine’s output is committed through Experience Builder: campaigns, journeys, channels, product experiences, rewards, and workflows. Delivery and performance data are written straight back, with no product boundary between decision and execution. A boundary would add a sync and a delay, and the signal the online learner needs would arrive after the next decision had already been made.

    The Feedback Loop

    Every delivered action produces one logged record:

    The outcome is written to TesseractDB™ as a campaign response, updates , and becomes part of both the context and the cap counts that define the feasible set for that user’s next request.

    What “live” means here is specific: once an outcome is observed, the next decision for that user is computed with it. The decision, the context it was made on, and the outcome it produced are never separated by more than one request cycle.

    Evaluating a Learner That Keeps Learning

    A bandit’s value is in how it adapts, so a single held-out accuracy number says little. Evaluation replays history in time order, in two phases:

    1. Frozen: score with the warm-started and no updates, to measure the starting policy.
    2. Learning: update from each observed outcome, as in production, to measure how quickly it adapts.

    Each replayed decision is posed as a choice among the observed action and a matched set of alternatives, constructed so the model is never asked to beat options the user was never eligible for. On that set we measure calibration and ranking quality, for example

    where is the set of decisions with a positive outcome. MRR is restricted to positives because a rank is only meaningful when the user acted; Brier runs over all decisions, non-responses included. Comparing the two phases shows directly whether online updates help, and how fast. Offline replay ranks candidate policies; the claim that a policy lifts the business goal is settled only against a randomized holdout.

    One engineering rule protects all of this: every model ships with a machine-readable contract describing exactly how its inputs are built, and serving builds requests from that contract. If serving prepared inputs differently from training, even by a monotone rescaling, learned weights would be applied on the wrong scale and the model would degrade silently.

    Trust, Security, Explainability, Governance

    In CleverAI™, trust, security, and governance do not sit at the end of the architecture as an approval box. They operate across it.

    Explainability falls out of the score’s structure. Because the online score is a sum of effects and interaction terms, every decision decomposes into the contribution of each user-side and action-side group, so a marketer can see why this user got this campaign rather than only that they did.

    Around that sit marketer-defined guardrails, frequency controls, RBAC, HITL approvals, full audit trails, and user-level inspection at every layer. This is what lets the system’s autonomy increase without the decision function becoming opaque.

    Technical Summary

    ComponentRoleMathematical objectMechanism
    Goal + Strategy + GuardrailsControl planeReward ; feasible set Objective and hard constraints, supplied to AI Studio and the decision function
    AI StudioAgentic automationCallers of Lifecycle, one-time, drop-off and custom agents; MCP Server & Skills; HITL and RBAC. Every agent draws its decision from the same policy
    TesseractDB™Real-time feature storeOne store for live events, 10-year event-level history, profile state, campaign responses, derived features
    Clean RoomFeature extension over non-exportable dataNon-reversible embeddings computed in the customer environment; split learning exchanges only activations and gradients
    CreatorsCandidate generation, Content variants; semantic embeddings plus structured metadata
    ExpertsScore providers, Recommendations, predictions, best time, best channel, sequence intelligence
    System 1: offline learnerUser prior, Campaign-agnostic non-linear model on leakage-free labels; warm-starts the online policy
    System 2: online learnerPer-user action selection, Contextual bandit with bilinear user × action terms; score-following exploration; per-outcome updates
    Experience BuilderExecutionCampaigns, journeys, channels, product experiences, rewards, workflows
    Feedback loopPolicy updateOutcome → TesseractDB™ → online learner → next decision, within one request cycle
    Trust & SecurityCross-cutting governanceAdditive score decompositionPer-decision explainability, RBAC, audit, HITL approvals, frequency controls

    Conclusion

    Live 1:1 personalization is not one model. It is a constrained decision problem, and every layer of CleverAI™ exists to keep one part of it correct at production scale: the context , the action description , the feasible set , the policy , and the reward that closes the loop.

    Five principles carry the design:

    1. Split intuition from deliberation. Who the user is changes over weeks; which option works changes with every launch. Separating System 1, and , from System 2, , lets each part be as heavy or as light as it needs to be.
    2. Let System 1 inform System 2. A rich offline representation gives a small online learner non-linear reach at almost no serving cost.
    3. Describe actions, don’t enumerate them. Scoring from , including semantic intent, lets learning transfer across campaigns and gives cold start a principled answer.
    4. Choose interactions on purpose. Personalization lives in the user × action terms, so deciding which ones the model learns is a first-class design decision.
    5. Constrain before you score. Guardrails as the support of the policy mean every output is valid, and the engine learns only from actions it actually chose.

    For a marketer, the result is a system that replaces the sequential test calendar with continuous, per-user experimentation, starts new campaigns from what similar ones already taught it, and stays inside the rules the business sets. For the AI Studio agents built on top, it means every agent draws on one policy and one memory of each user, so autonomy can grow without intelligence fragmenting across tools.

    Posted on October 6, 2026

    Author

    Jacob Joseph LinkedIn

    Heads Data Science.Expert in AI, Data & Analytics and awarded 40 under 40 Data Scientists in India.

    Please enter a valid work email

    Smiling Woman Holding Phone