DAI 4.1 · STEP-BY-STEP · learning/agents

Learning Agents & Sapphire

Define persistent learning behavior with observations, choices, rewards, imitation, dialogue, world grounding and knowledge import.

Build from scratchTest in Minecraft4.1 source-aligned

1. Create the agent file

The loader reads ordinary datapack JSON from data/<namespace>/learning/agents/. Learned values are stored outside the datapack by the client learning runtime, so editing the definition changes the allowed learning surface without embedding learned memory into the pack itself.

data/my_game/learning/agents/sapphire.json
{
  "name": "Sapphire",
  "enabled": true,
  "autonomy_default": false,
  "decision_interval": 10,
  "learning_rate": 0.15,
  "discount": 0.90,
  "exploration": 0.10,
  "success_reward": 0.25,
  "failure_reward": -0.50,
  "observations": [
    {"id":"hungry","condition":{"type":"player_hunger","operator":"<","number_value":8}}
  ],
  "choices": [
    {"id":"eat","action":"my_game:eat_food","demonstrations":["eat"]}
  ],
  "rewards": [
    {"id":"fed","condition":{"type":"player_hunger","operator":">=","number_value":18},"value":1.0}
  ],
  "imitation":{"enabled":true,"weight":0.5},
  "dialogue":{"enabled":true,"greeting":"Hello."},
  "grounding":{"enabled":true,"weight":1.0,"confidence":0.75},
  "knowledge":{"enabled":true,"learn_from_chat":true,"inference_enabled":true,"max_inference_depth":6,"import_folder":"dai_learning/knowledge_inbox","allow_grounding_imports":true}
}

2. Define what the agent can observe

Every observation wraps a normal DAI condition. That means the full 273-condition catalog is available as the sensory vocabulary where the condition makes sense.

3. Define choices

Each choice has an ID, an action reference and optional demonstration labels used for imitation. Keep the action library intentionally constrained to behaviors you want the agent to be able to execute.

4. Reward state transitions

Reward conditions fire on the false → true edge, rather than paying every tick while a condition remains true.

5. Dialogue, grounding and knowledge

SystemPurpose
DialogueLocal Sapphire conversation, response learning and continuation reward.
GroundingConnect ordinary language to the world object under the player's crosshair/raytrace.
KnowledgePersistent concepts/relationships, chat learning, bounded inference and safe JSON knowledge hotpatch imports.

Safety/design rule

Learning agents do not create new engine privileges. They choose only from actions and observations the definition/runtime already exposes.