Appendix B — 🪜 Cognitive Capacity: «Curate» 🙶Constructive Fill-in🙷

Tip

📊 Character Counts / 字數統計

  • Chinese version: 25482 characters
  • English version: 60122 characters

The ladder of capacity, spanning eons of time and space, living beings moving from survival to transcendence; its tiers advancing in turn are the course of the mind.

The union of capacity, mirroring an interactive cosmos, human and machine joining action into a duet; its combined dynamism is the echo of technology.

Elevating ❝mental fill-in❞ into 🙶constructive fill-in🙷, tracing the direction of “intent”, designing an intelligent system through “constructive scaffolding,” to weigh the “cost of intelligence” against the world’s energy.1

Continuing from the 🏄🏼verb «options» of the previous appendix, “💪Action”, and the 🌊noun «options» of “🧠Brain”, this appendix set, “🪜Cognitive Capacity”, aims to summarize and classify “🌊nouns ➕ 🏄🏼verbs” into “🏗️Constructive Scaffolding,” and, through the Double Diamond model, to demonstrate the “capacity-building” stages of an agentic system.

Tip B.1: 🪜 Cognitive Capacity Tip: «The Appendix Triad» “Action ~ Brain ~ Cognitive Capacity”
  • “💪Action”: Focusing on Verbs, the “Action Concerto”
  • “🧠Brain”: Focusing on Nouns, “Tao-Intelligence Practice”
  • ➜ “🪜Cognitive Capacity”: Focusing on Nouns ➕ Verbs, “Constructive Scaffolding”
    • 〜 Elevating ❝mental fill-in❞ into 🙶constructive fill-in🙷: the systemic-innovation capacity to reach a goal.

🪜 The core of human-machine decision-making is “curating” — 🙶constructive fill-in🙷 of the human-machine «choice», the systemic-innovation capacity to reach a goal: using the “🏗️Constructive Scaffolding” to diverge and produce combinations «options» of 🌊learned nouns ➕ 🏄🏼action verbs, so that the user and the AI, following intent, carefully curate («curate»), and converge from these into a key «choice».

This systematic “diverge first, then converge” of “«curating» and «choosing»” embodies the collaborative capacity of the “Action ~ Brain ~ Cognitive Capacity” triad that treats AI as a partner in shared action — choosing among options to make a «choice», achieving systemic dynamic orchestration, aligned to the goal.

💡🌌

🌊Noun ➕ 🏄🏼Verb🏗️Constructive Scaffolding” (Constructive Fill-in Scaffolding, or Constructive Scaffolding) emphasizes building cognitive capacity for the “unity of knowledge and action.”

To highlight the process of discernment involved in “acting where one should, and not acting where one should not” within “intent” and “will,” this book further calls the structure that lands in a specific situation, after careful curation, the “🧭Knowledge-Action Scaffolding” (Scaffolding Knowledge for Actions).

Applying this concept to the field of AI gives us “🤖AI Knowledge-Action Scaffolding” — the structure and context through which an agent effectively “participates” in the real world. This participation cannot be separated from the human mind’s symbol grounding of the intent and meaning behind action, and is chiefly manifest through:

When using the tables and figures in each appendix, readers may adopt a strategy of “diverge first in detail, then converge through careful curation”: first use 🏗️Constructive Scaffolding to expand «options», then use 🧭Knowledge-Action Scaffolding to “«curate» and «choose»” — a practical guide for expanding or trimming the system boundary on your own.

  1. Ontology Layer:
    • 🙶Constructive Fill-in🙷 vs. ❝Mental Fill-in❞: elevating passive guesswork (❝mental fill-in❞) into “a purposeful, cost-bearing, systematic construction (🙶constructive fill-in🙷).”
    • World Fragment: using “🏄🏼verb🌊noun” to compress and represent an operable, real situation.
  2. Method & Discernment Layer:
    • 🏗️Constructive Scaffold: logically expanding on 🌊nouns (knowledge) + 🏄🏼verbs (action) to maximize the generation of situational «options».
    • «Curate» (Selection & Curation): based on “intent” and “will,” carefully selecting from and paring down a large set of «options».
    • 🧭Knowledge-Action Scaffold / 🤖AI Knowledge-Action Scaffold (Action-Knowledge Scaffold): the concrete «choices» (Choices) that converge after “«curating»,” serving as the action structure through which human and machine jointly participate in the world.
  3. Mechanism Layer:
    • The SIPAEA cycle: Sense \(\to\) Frame \(\to\) Plan \(\to\) Act \(\to\) Evaluate \(\to\) Adapt, a dynamic progression.
    • The Mentor MoE: a multi-agent decision network based on a Mixture of Experts model.

🎓 This appendix corresponds to the methodological role within «The Appendix Triad» (for ontology and epistemology, see “🧠Brain”). For the complete tripartite comparison, see the academic-framework explanation in the Preface 〜 «The Appendix Triad».

🩵 The Proactive 🙶Constructive Fill-in🙷 🏗️

This appendix set, “🪜Cognitive Capacity”, holds that “Constructive Scaffolding” requires curating a «choice» in order to realize one’s own goals — a form of 🙶constructive fill-in🙷 (Constructive Fill-in) that involves “acting where one should, and not acting where one should not.”

  • Constructive fill-in: from ❝mental fill-in❞ to 🙶constructive fill-in🙷, the «options» of 🌊nouns ➕ 🏄🏼verbs.
  • 鷹架: refers to the innovative «choices» of knowledge and method that humans use when building AI.

Under the “Action ~ Brain ~ Cognitive Capacity” triad, then, systemic-innovation capacity is diverging among possible «options», systematically curating, and making a feasible «choice» based on the situation and need.

Note B.3: 🏗️ AI Knowledge-Action Scaffolding: The Constructive Scaffold Formula

🏗️ Constructive Scaffold 🟰 🌊Noun ➕ 🏄🏼Verb

🌊Noun➕🏄🏼Verb🟰Scaffold

To summarize and classify this “Constructive Scaffolding,” this book also calls it “AI Knowledge-Action Scaffolding,” to emphasize its operational definition of AI as one of “unity of knowledge and action” and “acting where one should”:

🏗️ AI Knowledge-Action Scaffolding 🟰 🌊Noun ➕ 🏄🏼Verb

To clearly express “what (noun) to do (verb) with AI,” this book recommends selecting «options» as needed and making a «choice»:

  • Nouns: as “knowledge” noun «options», divided into must-know, can-know, and not-yet-known (see the 🌊noun list);
  • Verbs: as “action” verb «options», divided into must-do, can-do, and need-not-do (see the 🏄🏼verb list);

When you can clearly express “what (noun) to do (verb) with AI,” it is like erecting scaffolding on a construction site — once the structure is up, you can build whatever result you need.

With a scaffold in place, readers can build or construct just about anything they set out to.

The capacity for systemic innovation lies chiefly in mastering divergent possibilities — «options» — and then converging purposefully on «choices». The concentric-circle system of nouns provided here, the 🌊noun list, and the verb 🏄🏼verb list, can systematically help us master «options» and make «choices».

😵‍💫 A Case Study: Expanding on Large Language Model Applications

Taking Large Language Models as an example, the discernment-and-innovation system of “Constructive Scaffolding” — 🌊nouns ➕ 🏄🏼verbs — can be laid out as follows:

  • 🔴 Core Circle (foundational scaffold):
    • 🌊 Nouns: corpus, data, analysis, decision algorithms, AI knowledge points, etc.
    • 🏄🏼 Verbs: survive, remember, understand, etc.
    • Constructive-scaffold examples:
      • “Survive + corpus; compute + energy” → the model relies on vast corpora and compute to keep running.
      • “Remember + data; compute + parameters” → parameter weights memorize vast patterns of text.
      • “Understand + semantics; embed + vector space” → contextual embedding vectors form semantic understanding.
  • 🟠 Middle Circle (interaction scaffold):
    • 🌊 Nouns: context, user intent, interactive tasks, community platforms
    • 🏄🏼 Verbs: belong, connect, control, apply
    • Constructive-scaffold examples:
      • “Belong + community platform” → the model is folded into a specific technology ecosystem, forming a sense of belonging through user habit.
      • “Connect + context” → interacting with humans through an API, establishing a mutual human-machine ❝mental fill-in❞ “connection” in a shared context.
      • “Control + interactive task” → in automated decision-making, the model assists the interaction flow, giving both user and platform a sense of control.
      • “Apply + user intent” → applying the model to generate answers, translations, summaries, and other user intents.
  • 🟢 Outer Circle (system scaffold):
    • 🌊 Nouns: world model, values, cross-domain knowledge, emergent theory
    • 🏄🏼 Verbs: apply/achieve, analyze/evaluate, contextualize/systematize, align/plan/integrate, create, construct/theorize/lead
    • Constructive-scaffold examples:
      • “Apply/achieve + professional domain” → completing professional tasks in medicine, education, law, and other fields.
      • “Analyze/evaluate + cross-domain knowledge” → attempting to analyze input and evaluate output, supporting cross-domain decisions.
      • “Contextualize/systematize + world model” → contextualizing information from different sources and attempting to generalize system knowledge (for the full discussion of the world model, see 🔖Appendix “🧠Brain”).
      • “Align/plan/integrate + values” → touching on the AI Alignment & Control Problem, aligning with human values and integrating diverse knowledge.
      • “Create + linguistic representation” → generating new text or narrative meaning, demonstrating creative potential.
      • “Construct/theorize/lead + emergent theory” → in research and application, forming new knowledge frameworks and even leading future modes of knowledge production.

As an appendix to this book, readers can bring themselves (or their students or children) to use systems thinking to contemplate useful “noun ➕ verb” combinations for self, environment, and world, and learn how to elevate ❝mental fill-in❞ into 🙶constructive fill-in🙷 through “Constructive Scaffolding.”

🌊Noun➕🏄🏼Verb Across Disciplines

This book’s AI Knowledge-Action Scaffolding is, in fact, a widely existing form of the “🌊noun➕🏄🏼verb” combination found across language and thought.

Its core value lies in using a concise semantic structure to compress “action” and “object” into an operable world fragment. This combination is not only a linguistic phenomenon — because of its functionality and creativity, it is applied across fields such as education, artificial intelligence, human-computer interaction design, and literary rhetoric, as shown in Table B.1.

Note B.4: 🌊Noun➕🏄🏼Verb: Operable World Fragments Across Disciplines
Table B.1: 🌊Noun➕🏄🏼Verb: Operable World Fragments Across Disciplines
Discipline 🏷️ Theory/Framework 📖 Description 🔤 Example
🎓 Education Bloom’s Taxonomy Learning objectives are typically phrased as “verb + noun,” verb = cognitive level, noun = knowledge domain analyze data, evaluate argument, create model
📚 Instructional Design Task-based Learning Tasks presented as “verb + noun,” clearly indicating action and object explain concept, design experiment, compare texts
🤖 Artificial Intelligence Task-oriented AI; Prompt Engineering Instructions typically use a “verb + noun” structure, matching human intuition and model parsing generate image, translate text, summarize article
💻 Human-Computer Interaction (HCI) Functional-grammar design Interface functions defined as “verb + noun,” for ease of operation and understanding click button, drag file, search data
🌐 Systems Thinking Modular action units In system modeling or curriculum design, verb + noun as an operable module build model, test hypothesis, optimize process
🗣️ Linguistics Functional Compounds Verb + noun combine to form a new word, directly describing a function or role lighter (打火機), pickpocket
🧠 Cognitive Linguistics Metaphor & Blending Verb + noun as “conceptual compression,” quickly generating a new image pick-me-up (提神飲料), dream-chaser
✍️ Literature / Poetics Rhetoric and metaphor-making Verb + noun combines to create a fresh image or character light-chaser, fire-thief, dream-chaser

In education, the learning objectives of Bloom’s Taxonomy are often phrased as “verb + noun” — for example, “analyze data” or “create a model” — where the verb marks the cognitive level and the noun defines the knowledge domain. “Task-based Learning” likewise relies on this combination to design learning activities, so learners can clearly understand “what to do” and “what to act upon.”

In AI and design, the “verb + noun” combination becomes the grammatical basis of human-machine interaction. AI has “task-oriented” forms; interface design often defines functions through structures like “click button” or “drag file,” while the Prompt Engineering behind Generative AI relies on instructions like “generate image” or “translate text” — matching human intuition while remaining easy for models to parse.

Continuing from ?tbl-actions-10 in the previous appendix’s verb «options», “💪Action”, the table below Note B.5 shows the operability produced once a noun is paired in, yielding “world fragments” of differing tiers and across differing domains.

Note B.5: 🌊Noun➕🏄🏼Verb: Operable World Fragments Across Disciplines
Bloom + Whole-Universe Ten Tiers⟩ 🌊Noun➕🏄🏼Verb: Cross-Disciplinary Examples {#nte-actions-10-nouns}
Tier 🧰 Verb 📊 Data/Information 🧪 Experiment/System 🤖 Machine/AI 📖 Text/Language 🤝 People/Society
10 🚀⚛ Construct/Theorize/Lead construct a data theory lead system design theorize an intelligent architecture construct a language model lead a social vision
9 🧭☯ Align/Plan/Integrate integrate data sources plan an experimental process align an AI module integrate a corpus plan social action
8 🪞🪟 Contextualize/Systematize contextualize data systematize experimental results contextualize AI output systematize text contextualize interpersonal interaction
7 🎯🛠️ Apply/Achieve apply a statistical method achieve an experimental goal apply an algorithm apply a rhetorical technique achieve an interpersonal agreement
6 📚🤓 Understand understand data patterns understand system operation understand AI reasoning understand semantic structure understand social norms
5 💾🤓 Remember remember data points remember experimental steps remember model parameters remember vocabulary and grammar remember interpersonal experience
4 🏘️👨‍👩‍👧‍👦 Belong belong to a data community belong to a research team belong to an AI ecosystem belong to a language community belong to a social group
3 💪☸ Control control a data flow control system operation control AI behavior control text generation control social power
2 🤝💞 Connect connect datasets connect experimental modules connect an AI network connect contexts connect a community
1 🥗🚰 Survive preserve data maintain a system keep an AI running sustain a language sustain a society

🔀 The Double Diamond Model

Having grasped the semantic unit of “AI Knowledge-Action Scaffolding” behind the “🌊noun➕🏄🏼verb” combination, using AI to solve a problem can be expected to mean finding and defining a generalized problem-and-solution pair of “verb + noun” — a pair that includes a continuous or cyclic “verb + noun” combination.

To systematically guide readers through the full journey from discovering a problem to designing a solution, this appendix adopts the well-known Double Diamond Model, dividing the design process into two “diverge” and “converge” cycles, used to define the final “verb + noun goal” and the related set of tasks.

The Double Diamond Model (abbreviated the Double Diamond) comprises four stages (Discover, Define, Develop, Deliver). It revolves around two core concepts — problem definition and solution — aiming to ensure that designers first “do the right thing” (solve the real problem), then “do the thing right” (deliver a good solution). This is one of the core tasks of an AI Product Manager.

Important B.1: The Double Diamond Model: Two Spaces of Problem and Solution
Table B.2: The Double Diamond Model: Two Spaces of Problem and Solution
Diamond Stages Core Purpose Space and Key Principle
♦️ Discover → Define Problem definition: ensuring
“we do the right thing.”
Problem space: through divergence and convergence, precisely focusing direction from broad insight.
💠 Develop → Deliver Solution: ensuring
“we do the thing right.”
Solution space: through divergence and convergence, choosing the best solution from many concepts.

The structure of these four stages ensures that before proposing any feasible solution, we first fully understand the problem, and only then precisely lock in a direction.

In building the capacity for developing intelligent solutions, the double-diamond structure offers effective guidance: the first diamond helps us diverge to explore the problem and converge on a goal; the second diamond helps us diverge to develop solutions and converge on a deliverable product.

The table below shows how this appendix’s sections correspond to the four stages of the Double Diamond, ensuring the setting and execution of an agent’s goals both have a clear method to follow:

Stage Corresponding Section Title Core Content
Discover 🏗️ Exploring Noun + Verb Combination Goals Fully validating problems and opportunities. Systematically diverging from the three noun categories and three verb categories to surface every possible intelligent combination that can achieve a “goal” (such as solving a pain point or achieving a growth opportunity). Clarifying the options and limits of individual, machine, and human-machine collaboration with respect to the goal.
Define 🎯 Locking In the Problem 🙶Goal🙷: Defining a Career-Assistant Case Study 🛡️ Using question-formulation methods like HMW (How Might We) to select an operable “noun + verb” grammatical scaffold, converging the broad results of exploration into a single, actionable intelligent-goal definition.
Develop 🌌 Building Capacity Solutions Level by Level Based on the defined goal, developing different agent architectures and toolchain prototypes (e.g., adopting RAG, multi-agent collaboration, etc.). Applying the “verb flow/cycle” scaffold options, combined with systemic innovation and capacity-building, to generate prototypes.
Deliver 🪜 Choosing a Solution 🙶Constructive Fill-in🙷: Landing the Career-Assistant Case Study Converging from many candidate solutions onto a deployable Minimum Viable Product (MVP). Testing and validating the defined goal metrics, iterating the solution prototype, ensuring it lands reliably and accurately and delivers value.

In short, this appendix gathers the complete thought process required for an agentic system’s goal-setting, architecture exploration, and capacity-building, using the combination of “noun + verb” to make abstract needs and intelligent capacities concrete. The practical framework summarized here aims to help readers diverge a large set of “options” and transform them into «choices» with clear intent.


1. 🏗️ Exploring Noun+Verb Combinations

This section will use the “noun + verb” grammatical scaffold to systematically diverge every intelligent goal that could correspond to “solving a pain point” or “achieving growth,” and validate the feasibility and potential value of each combination. The point here is not to immediately find the one right answer, but first to broadly generate “options,” and then gradually converge onto a “choice.”

🌊 Noun «Options»

Nouns represent knowledge points or objects of information, and can be divided into three categories:

  • Must-know: core information or objects that must be grasped, such as “student profile,” “career path,” “skill requirements,” “further-education plan.”
  • Can-know: information that can be obtained through effort or tools, such as “industry trends,” “mentor resources,” “learning platforms.”
  • Not-yet-known: information that is currently impossible to fully grasp or highly uncertain, such as “the future shape of the workplace,” “emerging occupations,” “AI’s long-term impact on employment.”

This three-way distinction helps a designer judge which information needs to be filled in immediately, which can be gradually explored through an AI system, and which needs to be left flexible to accommodate uncertainty.

🏄🏼 Verb «Options»

Verbs represent modes of action, and can be divided into three categories:

  • Must-do: core actions that must be completed, such as “explore,” “analyze,” “plan.”
  • Can-do: supporting actions that may optionally be completed, such as “compare,” “simulate,” “connect.”
  • Need-not-do: actions that, in a given context, can be temporarily set aside, such as “over-prediction,” “redundant verification.”

This classification helps a team, when designing an intelligent goal, focus on the actions that truly create value, avoiding “too many actions” or “wasted resources.”

🧩 Grammatical Combination: What Should AI Do

When you can clearly express “what (noun) to do (verb) with AI,” it is like erecting scaffolding on a construction site — once the structure is up, you can build the result you need. For example:

  • “Use AI to analyze (verb) skill requirements (noun)”
  • “Use AI to plan (verb) a learning program (noun)”
  • “Use AI to connect (verb) mentor resources (noun)”

These grammatical combinations are not merely a way of speaking — they are a process of turning abstract needs into concrete design tasks.

🔎 Exploring Goal-Statement Options

Through the “noun + verb” semantic frame, a perceived exploration goal can be precisely expressed in terms of a concrete action and an object. Each such expression represents a path by which humans cultivate self-awareness and cognitive capacity across the material world, the social world, and the technological world.

Through these eight tiers of “verb–noun” pairings, readers can gradually build the capacity and portfolio of items needed to live in an “intelligent society.” These “exploratory combinations” will, in the next section, be further converged, through the HMW questioning method, into a single, actionable intelligent goal.


2. 🎯 Locking In the Problem 🙶Goal🙷: Defining a Career-Assistant Case Study 🛡️

Having listed the “problem combinations,” we can more clearly define the goal, thereby determining the choices and limits of individual, machine, and human-machine collaboration. Clearly defining the real problem is what determines the direction of any subsequent solution design.

This section demonstrates, through a case study, how to use question-formulation methods such as HMW (How Might We) to converge the broad results of exploration into a single, actionable intelligent goal definition, clearly setting out the choices and limits of individual, machine, and human-machine collaboration in a career-assistant scenario.

🧍‍♂️ The Individual’s Choices and Limits

The human capacity for ❝mental fill-in❞ comes from experience, situation, and imagination, allowing speculation and creation when information is incomplete. But this capacity is also constrained by cognitive bias (such as overconfidence, selective attention), emotional energy (fatigue and anxiety reduce judgment), and the boundary of knowledge (difficulty crossing disciplinary boundaries without outside information).

So while individuals can propose creative ideas in career planning, they often need outside support to fill in their blind spots. To overcome these limits, and increase the options for individuals to help and understand one another, the following design methods can be used to understand the needs of stakeholder individuals, in service of clarifying the “goal”:

  • 👤 Persona-Based Role-Enactment: designers can build personas such as “the exploring high-schooler,” “the anxious parent,” “the career counselor,” “the study-abroad consultant,” or “the trade-craft mentor,” and act out these roles in a workshop to simulate their language, emotion, and needs when making career choices. This can reveal an individual’s “mental fill-in” strategies and cognitive biases when information is lacking.

  • 🔮 Future Role-Play: having a student or designer imagine “myself five years from now” interacting with an AI assistant can surface an individual’s anxiety and expectation about future uncertainty, and help the designer understand a student’s needs across different time scales.

  • 🔄 Stakeholder Mapping + Role Rotation: designers take turns playing student, parent, and expert, experiencing the tension of each position, helping them understand an individual’s limits when making decisions, and revealing the conflicts and complementarities between roles.

With this concrete information and understanding in hand, one can clearly define a “goal” that aligns with stakeholder needs and expectations, designing a more targeted intelligent-assistance solution in light of the individual’s choices and limits.

🤖 The Machine’s 🙶Constructive Fill-in🙷 and Its Cost

The machine’s capacity for “fill-in” comes from massive data and compute, letting it quickly generate answers, simulate many possibilities, and outperform humans at large-scale information processing. But the cost of this capacity is: energy consumption (large-model computation), ethical risk (biased data), and a lack of situational awareness (an inability to truly understand human values and emotion). So while machine fill-in is highly efficient, it needs human oversight and adjustment.

Although AI can, through Large Language Models, quickly fill in information and offer suggestions, a designer still needs to turn its output into an actionable “goal”:

  • 📝 Prompt Templates: turning persona output into role-based prompts, letting the system adjust tone and content to different roles (e.g., “As a high-schooler, please explain… in simple language”).
  • 📈 Long-Term Scenario Simulation: turning insights from future role-play into retrieval and generation strategies — a RAG system can provide simulated responses for “workplace trends five years from now” or “future skill needs.”
  • 🌐 Multi-Perspective Retrieval: using the results of role rotation to design multi-perspective answers, where the same question is presented from a “student’s perspective,” a “parent’s perspective,” and an “expert’s perspective,” helping the user compare and balance different views.

With this concrete information in hand, a designer can clearly define a “goal,” using the retrieval and generation capacity of Large Language Models to lock in the problem 🙶goal🙷, and ensure the output can be correctly understood and adopted by humans.

🤝 The Human-Machine Scaffold

The key to human-machine collaboration lies in orchestrating two different forms of fill-in: the human “mental fill-in” supplies value judgment and situational understanding; the machine’s “fill-in” supplies data breadth and computational depth. When the two work together within a grammatical framework (verb + noun combination), a “duet” emerges: humans set the goal and ethical boundary, machines generate solutions and iterate quickly, and humans then filter and revise. This mode both avoids the blind spots of either side alone, and improves the comprehensiveness and feasibility of decisions.

Based on the “goal” options and limits from the previous stage of the double-diamond model, this stage begins to converge, locking in the problem 🙶goal🙷. Question-formulation methods such as HMW (How Might We) can be used to converge the broad results of exploration into a single, actionable intelligent-goal definition.

Taking the career assistant as an example, its core question — “How might we help students and parents…?” — can be concretely converged into the following actionable “noun + verb” grammatical frames:

  • 🎯 Explore Career Pathways
  • 📊 Analyze Skill Demands
  • 🧭 Plan Learning Programs and Portfolios
  • 🤝 Connect Mentorship Resources
  • 🧪 Test Decision Assumptions

Through this convergence process, the development team for an intelligent career assistant can, based on its own capacity and market intelligence, select the single most core, most worth-solving problem 🙶goal🙷, and gain a clearer grasp of the corresponding pain point to solve or growth opportunity to achieve.


3. 🌌 Building Capacity Solutions Level by Level

Based on the locked-in goal, this section will unfold, in stages, multiple agent architectures or toolchain prototypes (such as RAG, multi-agent collaboration, etc.), combined with a “verb flow/cycle” scaffold, to gradually build, test, and optimize a feasible solution.

This appendix set continues with the example, “How might we build a career assistant for high-school students?”, to demonstrate how to build an architectural solution, targeting the following “noun + verb” combination as the locked-in goal:

🧭 “How might we help students and parents… plan a learning program and pathway?”

🔄 The Flow/Cycle Scaffold

Once familiar with the “noun➕verb” combinatorial grammar, the next stage is to build the system level by level. This appendix introduces a verb flow/cycle scaffold template to help draft a concrete, dynamically capable cycle.

This SIPAEA scaffold template integrates the Sense–Plan–Act architecture from cybernetics and robotics with the Sense–Interpret–Adapt Flywheel model from organizational learning and adaptive systems:

🛰 Sense → 🪟 Frame/Interpret → 🗺️ Plan → 💪 Act → 🧮 Evaluate → 🔂 Adapt

Designers and builders can adjust, add, remove, or replace verbs as needed, to build the innovation flow best suited to their own situation. For instance, feedback can be added as a meaningful return-and-correct loop between any two actions.

The knowledge sources for this scaffold template are as follows, explained further in the annotated bibliography that follows:

  • Cybernetics and systems theory (Wiener, Beer) supply the engineering-mathematical and practical foundation for adaptation and feedback.
  • Decision science and military theory (Boyd, Endsley) supply dynamic models of perception, interpretation, and action.
  • Management and learning theory (Deming, Kolb) supply cycles of continuous improvement and reflection.
  • AI architecture (Russell & Norvig) implements these cyclic patterns in the design of intelligent agents, through the basic “sense–plan–act” architecture, extended further with “evaluate–adapt,” letting a system continually learn and optimize in a complex environment.
Table B.3: 🔄 The SIPAEA Scaffold Template Mapping Table
Source/Theory 🛰Sense 🪟Frame/
Interpret
🗺️Plan 💪Act 🧮Evaluate 🔂Adapt
Cybernetics and Systems Theory Observe/Monitor (Wiener) Learn/
Correct (Beer)
AI and Robotics Architecture Decide/
Conceptualize (Newell
& Simon)
Execute/Operate (Brooks, Russell & Norvig)
Decision Science and Military Theory Observe (Boyd), Monitor (Holling) Orient/
Reflect
(Endsley)
Learning and Management Cycles Monitor (Holling) Reflect
(Kolb)
Plan (Deming), Conceptualize (Kolb) Experiment/Act (Kolb) Check (Deming), Evaluate (Holling) Correct (Kolb), Adapt (Holling)

📚 Annotated Bibliography

The following are the specific sources for this book’s proposed scaffold template, and their corresponding verb sets:

  • ⚙️ Cybernetics and Systems Theory
    • Norbert Wiener, Cybernetics (1948) – lays the foundation of cybernetics, emphasizing the cycle “Sense \(\to\) Feedback (Evaluate) \(\to\) Adapt.”
    • Stafford Beer, Brain of the Firm (1972) – the Viable System Model: a recursive “Sense \(\to\) Interpret \(\to\) Plan \(\to\) Act \(\to\) Adapt” that ensures an organization’s survival.
  • 🤖 AI and Robotics Architecture
    • Newell & Simon (1972), Human Problem Solving – an early cognitive architecture, proposing the SIPA cycle of “Sense \(\to\) Interpret \(\to\) Plan \(\to\) Act.”
    • Russell & Norvig, Artificial Intelligence: A Modern Approach (3rd ed., 2010) – describes the Sense \(\to\) Plan \(\to\) Act architecture, further extended to “Evaluate \(\to\) Adapt.”
    • Brooks, subsumption architecture in robotics (1986, 1991) – a layered “Sense \(\to\) Act” loop, with “Adapt” capability.
  • 🎖️ Decision Science and Military Theory
    • John Boyd’s OODA loop (1987) – “Observe (Sense) \(\to\) Orient (Interpret) \(\to\) Decide (Plan) \(\to\) Act,” widely used in strategy and decision science; modern management literature typically supplements this with “Evaluate \(\to\) Adapt.”
    • Mica R. Endsley, Toward a Theory of Situation Awareness (1995) – proposes three tiers of situational awareness: “Perceive (Sense) \(\to\) Comprehend (Interpret) \(\to\) Project (Plan),” corresponding to the Sense–Plan–Act architecture of an intelligent agent.
  • 🔂 Learning and Management Cycles
    • Deming’s PDCA cycle (Plan–Do–Check–Act, 1950s) – an iterative improvement cycle, “Plan \(\to\) Do \(\to\) Check \(\to\) Act,” directly corresponding to “Plan \(\to\) Act \(\to\) Evaluate \(\to\) Adapt/Correct.”
    • Kolb’s Experiential Learning (1984) – a learning cycle: “Concrete Experience (Sense) \(\to\) Reflective Observation (Interpret) \(\to\) Abstract Conceptualization (Plan) \(\to\) Active Experimentation (Act).”
    • Adaptive Management (Holling, Walters, 1978–1986) – ecological management: “Monitor (Sense) \(\to\) Interpret (Interpret) \(\to\) Plan (Plan) \(\to\) Act (Act) \(\to\) Evaluate (Evaluate) \(\to\) Adapt (Adapt).”

🌉 LLM AI Engineering

Following the scaffold template of the verb flow/cycle, we can concretely break down how to achieve “planning a learning program and pathway for students and parents…”, as shown in the table below:

Verb Step Constructive Capacity (AI Knowledge-Action Scaffolding) Task Goal Engineering Implementation and Technical Detail
🛰 Sense Record needs and the student profile Capture comprehensive user and environmental data Multimodal API or MCP: extract a transcript (structured PDF extraction) or a portfolio (image/text embedding) uploaded by the user. Use Named Entity Recognition (NER) to extract key knowledge points.
🪟 Frame/
Interpret
Understand, contextualize Turn data into actionable situation and intent Context Engineering and role-based Prompt Engineering: give the LLM the role of “career counselor,” using RAG to retrieve the success paths of similar students (case studies), for situational framing.
🗺️ Plan Plan, apply Generate staged, diverse action paths Chain-of-Thought (CoT) or a planning agent: have the LLM first decompose the goal internally. Call external tools (e.g., a Course API) to gather parent or student discussion.
💪 Act Connect, achieve Execute the external operations of the plan Tool Use and external API connection: the LLM outputs a structured JSON request; an executor calls a job-platform (e.g., LinkedIn) API to retrieve job openings, or a course-platform API to retrieve course links.
🧮 Evaluate Evaluate, control Measure the effect of actions and user feedback LLM as evaluator and RLHF data collection: process user feedback and convert it into a quantified reward signal. Use agent evaluation to check API response speed and link validity.
🔂 Adapt Adapt, align Optimize the model and process based on evaluation Context Engineering supervised fine-tuning: use high-quality “success examples” and “failure examples” to improve the agent, ensuring the next planning pass better matches the user’s values.

In this way, achieving a “learning program and pathway” useful to parents and students requires a verb flow/cycle scaffold template, one that can also systematically incorporate services — like the job market — that change over time and context.

🧭 Multi-Agent Architecture and Prototype Connection

Having listed the various “verb flow/cycle” scaffold options, the next step is to progressively build, test, and optimize a feasible solution. The demonstration problem: “How might we plan a learning program and pathway for high-school students and their parents?”

🪟🧭 Multi-Agent Architecture

Multi-agent architecture: precise division of labor and multi-dimensional alignment. In the task of planning a learning program, a multi-agent architecture ensures the learning path is not just academically feasible, but also accounts for the resources and risks parents care about (e.g., time cost, return on investment). An MoE Router Agent routes “course selection” questions to an academic expert, and “cost” questions to a financial advisor.

The core mechanism of the MoE is to decompose the “career coach” role into several expert agents (e.g., an academic curriculum expert, a workplace-skills mentor, a family financial advisor). These experts, specialized in “learning-path planning,” together form a Mentor MoE (Mixture of Experts) using large language models and databases.

Regarding “life opportunity” expert roles, a Mentor MoE is formed using large language models and databases. This intelligent system can provide decision support based on the student’s profile, following the steps “🛰Sense \(\to\) 🪟Frame \(\to\) 🗺️Plan \(\to\) 💪Act \(\to\) 🧮Evaluate”:

  • 🛰 Sense: call relevant university-program data, job-market data, and current-and-projected data through API or MCP, so the Mentor MoE obtains data appropriate to the student’s specific time and place.
  • 🪟 Frame: use Prompt Engineering to design and, based on the time-and-place data, precisely frame the role positioning of the mentor expert group MoE — e.g., a university-and-major advisor for a given country, or a human-resources agent for a given city — helping to frame the dialogue that interprets the relevant data.
  • 🗺️ Plan: next, generate a roadmap connecting the training program to the subsequent career path, in particular using Retrieval-Augmented Generation (RAG) to expand into a real knowledge base (such as from the professional social network LinkedIn), providing a variety of roadmap options.
  • 🔨 Act: use Context Engineering to stitch together the student’s personal history of dialogue with the mentor expert group MoE and newly received information, ensuring coherence.
  • ✅ Evaluate: apply the Agent Reliability & Evaluation knowledge point to challenge a temporary claim made by the student or mentor expert group, embracing uncertainty, offering multiple alternative paths such as a “pivot,” and appropriately factoring in additional considerations and evaluations of life opportunity and risk that the student and parents care about — such as the chance of finding a partner, the expected timeline for starting a family, or the city and lifestyle associated with a given career.

Through this five-step cycle, the MoE Career Mentor Network can dynamically coordinate across complex scenarios, retaining deep domain insight while still integrating into a coherent overall recommendation.

🔗📒 Toolchain Prototype

Core value: efficient execution and transparent tracking. The value of a toolchain lies in turning a planned learning path into an executable, trackable task list, and using collaborative filtering to improve the efficiency of resource recommendation.

Connecting the matching resources of various real-world domains with platform-side Collaborative Filtering tools forms a toolset supporting “life opportunity” decisions. This toolchain focuses on various matching services, helping students explore potential career-option matches, major matches, partner matches, city or country matches, and so on, following the steps “🛰Sense \(\to\) 🪟Frame \(\to\) 🗺️Plan \(\to\) 💪Act \(\to\) 🧮Evaluate” based on the student’s profile:

  • 🛰 Sense: through an initial questionnaire and dialogue, use API or MCP to structurally extract the student’s:
    • core values, interests, and personality traits (VIP), along with expectations for future life (e.g., expected salary, work-life balance, timeline for starting a family), as the base data for all subsequent matching;
    • learning style, current skill level, and preferred content format (e.g., video, text, hands-on practice), as the base data for all subsequent matching.
  • 🪟 Frame: the core here is combining Collaborative Filtering technology to compare the student’s VIP data against diverse data sources from government, industry, and academia, framing and recommending learning resources that best match the student’s learning preferences and knowledge gaps — such as senior students or career mentors with similar traits or comparable success paths.
  • 🗺️ Plan: run multiple independent matching algorithms, and use Retrieval-Augmented Generation (RAG) to expand into:
    • a knowledge base of course reviews and prerequisite-skill requirements, offering multiple combinatorial learning-path options, computing and weighting each option’s completion and efficiency score;
    • a matching knowledge base (e.g., city cost of living, immigration policy), offering multiple combinatorial life-path options, computing and weighting each option’s fit score.
  • 🔨 Act: use Prompt Engineering to build an interactive module letting students pose hypotheses or test scenarios against matching results or a specific path, and then use Context Engineering to update matching weights and parameters in real time, optimizing the algorithm’s results.
  • ✅ Evaluate: apply the evaluation framework of Agent Reliability & Evaluation to challenge the student’s current choice or current learning progress, surfacing the potential risk and life opportunity of each path (such as missing an application deadline, industry longevity, or the chance of finding a partner), ensuring the student makes their final progress-management choice with a full understanding of execution uncertainty.

This toolchain example shows how to modularly integrate multiple services, using a loop to ensure each recommendation more closely matches user needs.


The two prototype architectures built up step by step above concretely show how a “verb flow/cycle” scaffold can develop solutions of different natures. Worth noting:

  • 🔂 The engineering necessity of adaptation:
    • This stage is not merely a conceptual “correction” — in engineering terms, it requires the system to have continuous integration and deployment (CI/CD) capability, using the user feedback collected in the 🧮Evaluate stage for model fine-tuning or automatic knowledge-base updates.
    • For fast-changing exam or academic standards, adaptation is the key mechanism ensuring data freshness. To pursue relative stability of a career, one instead needs annually updated data to confirm the relative stability of the path — this is itself a dynamically stable act of adaptation.
  • 🪟 The decision value of framing:
    • In the LLM era, framing is a pivotal process linking past to future. It shows concretely how, within the “solution divergence” stage, one carries out data divergence and convergence for a specific user.
    • The quality of framing directly determines the subsequent relevance and bias risk of the LLM’s output. Precise framing through context engineering and prompt engineering helps the LLM better perceive the real world and the process of discovering personal pursuits.

These two solutions demonstrate how, in the Develop stage of the second diamond, a six-step verb scaffold can diverge into deliverable solution options, from which the Minimum Viable Product (MVP) best able to validate the need can be refined for delivery, while accumulating generalizable patterns reusable in future scenarios.

4. 🪜 Choosing a Solution 🙶Constructive Fill-in🙷: Landing the Learning Program 🛡️

In the final stage of the double-diamond model, Deliver, our goal is to converge, from the many prototypes (multi-agent or toolchain) produced in the “Develop” stage, onto the Minimum Viable Product (MVP) best able to validate the core value. This stage emphasizes testing, validating metrics, and continuous iteration in a real environment, ensuring the intelligent learning solution can be reliably deployed and continue creating value.

📌 Converging and Validating the MVP

To successfully “deliver” a useful learning-program system, one must combine the thinking behind Agent Reliability & Evaluation and the AI Product Manager. When building the Minimum Viable Product (MVP), we focus on helping student and parent users “plan a learning program and pathway,” in order to test the solution’s timeline feasibility, resource fit, and value alignment.

🪟🧭 Multi-Agent Architecture MVP

The multi-agent architecture solution is characterized by layered planning and expert-resource alignment:

  • Goal: to validate the effectiveness and resource fit of a layered expert group combined with the core process, within the multi-agent architecture.
  • MVP content: two sub-agents — the “academic curriculum expert” and the “family financial advisor” — are selected as the core of the MVP. The system uses RAG to retrieve the latest exam syllabus.
    • Academic expert: focused on generating a staged learning-content list and timeline plan.
    • Financial advisor: focused on cost estimation and risk flagging for the external resources in the learning-content list (such as tutoring or online-course fees), based on the parent-provided budget range.
    • This collaborative mode ensures the delivered learning solution is not only “academically feasible” but also “financially feasible.”
  • Delivery and monitoring: a back-end dashboard monitors the plan’s “budget fit” (validating the financial advisor’s effectiveness) and “timeline completion rate” (validating the academic plan’s effectiveness). We need to ensure response quality (e.g., via Agent Reliability & Evaluation) meets expectations.

🔗📒 Toolchain MVP

The toolchain solution focuses on coordinating and outputting learning resources:

  • Goal: to validate the efficiency and precision of the toolchain mode on a specific, high-frequency task of matching learning resources and producing outputs.
  • MVP content: launch a combination of “style-matching algorithm + structured output,” specifically solving the tasks of “course matching” and “portfolio optimization.”
    • Course matching: using a Collaborative Filtering algorithm, based on the student’s learning style (e.g., visual, hands-on) and current knowledge gaps, to recommend the best-fit online courses, books, or hands-on projects.
    • Output optimization: applying Prompt Engineering and Context Engineering, letting a user input a piece of learning output (such as a project summary), which the system optimizes into a professional portfolio narrative that meets university-application standards, exported as a structured JSON/PDF format.
  • Delivery and monitoring: the user sees a “course fit score” in the UI. The system applies Prompt Engineering to let students or parents update the student’s situation, and displays dashboard data in the UI, letting users iteratively update prompts and generation results from the perspective of an AI Product Manager, measuring “structured-output acceptance” — i.e., the number of times a user edits the system’s generated portfolio description — to confirm whether the solution effectively helps users diverge and converge on【learning-resource options】.

🏋👨‍🏭👨🏽‍⚕️ Collaborative-Network MVP

Combining the two solutions above, and factoring in parent–student–agent collaboration and value alignment at a further level, we build an integrated collaborative-network MVP:

  • Goal: to validate whether combining agent interaction with parents can help align both parties’ learning goals and values, reducing decision conflict.
  • MVP content: integrating a “MoE learning plan” and “values dialogue” module.
    • The system first gives students and parents separate independent values questionnaires (e.g., valuing stability vs. high-earning potential), and then has the MoE integrate the results.
    • The agent surfaces value differences between the two sides during dialogue, and offers a dialogue framework for bridging disagreement — for example: “The academic expert suggests Path A, but the parent leans toward Path B; let’s simulate the long-term impact of each choice.”
  • Delivery and monitoring: the system uses Agent Reliability & Evaluation to monitor a “dialogue conflict index” (e.g., density of emotionally charged language) and a “goal-consensus rate,” ensuring the network becomes a safe “collaborative decision-making” environment that can withstand pressure from any single family’s or society’s values.

These three MVP examples demonstrate how, after divergence in the Develop stage, organically combining knowledge points like Retrieval-Augmented Generation (RAG), Prompt Engineering, Context Engineering, Agent Reliability & Evaluation, and the AI Product Manager lets us quickly converge on the minimum viable product that best validates the value of learning-program planning.

🚀⚛ Engineering “Fill-in”

Looking back over this appendix set, we started from the core concept of “verb flow/cycle,” making the abstract notions of capacity and mind concrete into an executable, collaborative action scaffold, successfully realizing an engineering closed loop from divergent «options» to focused «choices».

Here, the noun «choice» is used to draw the whole appendix together — from theoretical framework, through process breakdown, to MVP delivery — becoming the basic methodology for building an AI intelligent system.

We built an action blueprint through the “verb flow/cycle scaffold”:

  • ⚙️ Core value: modularizing and tracking the decision process (from Sense to Adapt).
  • 🔄 Key cycle: focused on Frame/Interpret and Evaluate, to ensure relevance of input and alignment of output.
  • 🤝 Human-machine co-construction: folding an LLM into the process, upgrading the machine from a mere tool into a collaborative partner capable of sensing, interpreting, planning, acting, evaluating, and adapting.

We further landed this theoretical architecture into agent prototypes:

  • 🪟🧭 The MoE collaborative engine: embodying “professional division of labor and alignment.” Through the layered collaboration of the academic expert and the financial advisor, resolving the dual considerations of academic feasibility and financial feasibility when planning a learning program.
  • 🔗📒 Toolchain-strengthened matching: embodying “efficient execution and transparent tracking.” Through collaborative filtering and structured output, turning a planned path into an executable, trackable task list, and giving the LLM genuine external-action capability.

Returning to our opening question: how do we achieve a “learning-program planning” useful to parents and students? How do we turn a verb flow into engineering practice?

  • 🪜 A converging constructive scaffold: following the principle of “diverge first in detail, then integrate,” converging scattered «options» into an ordered «choice», elevating the human-machine ❝mental fill-in❞ into a verifiable, transferable 🙶constructive fill-in🙷. Through the design of agents and toolchains, aligning “learning resources,” “budget,” and “target values.”
  • 🎯 From aligned intent to design: presenting one’s own “direction, boundary, and standard” through goal-oriented discernment. Through a value-alignment MVP, challenging student and parent perceptions of “value differences,” letting capacity truly land as a systematic outcome that meets shared expectations.

In short, because machine learning in AI is still unable to autonomously determine and revise its “objective function,” the “goal” of the human as user and designer is to consciously align learning planning with values, in response to the needs of students and parents. In this process, the grammar of “noun + verb” combinations converges into a solution «choice», forming the complete engineered system running from action, to capacity, and on to mind.

A three-part duet: with “Action ~ Brain ~ Cognitive Capacity” as its main melody, treating AI as a partner in shared action, letting 🙶constructive fill-in🙷 be not merely a patch, but an innovation that gives rise to new dynamics.

Footnotes


  1. What this passage is trying to express is a scientific-evolutionary perspective on contemporary AI:

    • The ladder of capacity: refers to the process of mental growth across long evolution, gradually moving from the most basic pursuit of “survival” to the pursuit of “transcendence.”
    • The union of capacity: emphasizes that human and machine collaborate within an interactive cosmos, and that the duet of their actions displays the power and echo that technology brings.
    • Seen from the world’s perspective: from the imaginative guesswork of the human brain (mental fill-in), to the capacity of AI (auto-completion), we must ask about the intent and will behind them, using “constructive scaffolding” to guide the development of intelligent systems, while also considering whether the energy cost is worth it.
    ↩︎