{
  "base_url": "http://127.0.0.1:18000/v1",
  "budget": {
    "max_completion_tokens": 9216,
    "reasoning_headroom_tokens": 8192,
    "visible_answer_tokens": 1024
  },
  "checks": [
    "smoke",
    "json",
    "needle",
    "tools",
    "streaming-tools",
    "tool-result"
  ],
  "evidence_policy": {
    "allowed_finish_reasons": [
      "stop",
      "tool_calls"
    ],
    "errors": [],
    "reasoning": "any"
  },
  "image_expect": [],
  "model": "llm.primary",
  "multimodal_input": null,
  "observations": [
    {
      "content": "```python\nsum(xs)\n```\n\nThat's it \u2014 Python's built-in `sum()` does exactly this. If you need it as a function:\n\n```python\nget_sum = lambda xs: sum(xs)\n```\n\nAnd if you want to do it manually (e.g., for learning purposes):\n\n```python\nfunctools.reduce(lambda a, b: a + b, xs, 0)\n```",
      "content_chars": 278,
      "content_excerpt": "```python\nsum(xs)\n```\n\nThat's it \u2014 Python's built-in `sum()` does exactly this. If you need it as a function:\n\n```python\nget_sum = lambda xs: sum(xs)\n```\n\nAnd if you want to do it manually (e.g., for ",
      "finish_reason": "stop",
      "passed": true,
      "reasoning_chars": 711,
      "reasoning_excerpt": "The user wants a Python one-liner that returns the sum of a list `xs`. This is a very simple request. The most obvious answer is:\n\n```python\nsum(xs)\n```\n\nThat's the idiomatic Python way. There's also ",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "seconds": 4.074,
      "test": "smoke",
      "usage": {
        "completion_tokens": 287,
        "prompt_tokens": 28,
        "total_tokens": 315
      },
      "validation_detail": "contains sum("
    },
    {
      "content": "{\"language\":\"python\",\"ok\":true}",
      "content_chars": 31,
      "content_excerpt": "{\"language\":\"python\",\"ok\":true}",
      "finish_reason": "stop",
      "passed": true,
      "reasoning_chars": 187,
      "reasoning_excerpt": "The user wants me to return ONLY a JSON object: {\"language\":\"python\",\"ok\":true}. No prose.\n\nThis is a simple instruction. I should output exactly that JSON object with no additional text.",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "seconds": 0.797,
      "test": "json",
      "usage": {
        "completion_tokens": 52,
        "prompt_tokens": 30,
        "total_tokens": 82
      },
      "validation_detail": "parsed keys=['language', 'ok']"
    },
    {
      "content": "ZEBRA-42917-QUARTZ",
      "content_chars": 18,
      "content_excerpt": "ZEBRA-42917-QUARTZ",
      "finish_reason": "stop",
      "passed": true,
      "reasoning_chars": 559,
      "reasoning_excerpt": "The user is asking me to extract a launch code from a large block of repetitive text. The text contains an embedded message that says \"IMPORTANT: The launch code is ZEBRA-42917-QUARTZ.\"\n\nThis appears ",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "seconds": 3.375,
      "test": "needle",
      "usage": {
        "completion_tokens": 134,
        "prompt_tokens": 6803,
        "total_tokens": 6937
      },
      "validation_detail": "needle present"
    },
    {
      "content": "",
      "content_chars": 0,
      "content_excerpt": "",
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 78,
      "reasoning_excerpt": "The user wants the weather in Oakland. I have a get_weather tool. Simple call.",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "request_index": 3,
      "seconds": 4.908,
      "test": "tools",
      "usage": {
        "completion_tokens": 31,
        "prompt_tokens": 4967,
        "total_tokens": 4998
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content": "",
      "content_chars": 0,
      "content_excerpt": "",
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 81,
      "reasoning_excerpt": "The user wants the weather in Oakland. I have a get_weather tool. Let me call it.",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "request_index": 0,
      "seconds": 4.909,
      "test": "tools",
      "usage": {
        "completion_tokens": 33,
        "prompt_tokens": 4967,
        "total_tokens": 5000
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content": "",
      "content_chars": 0,
      "content_excerpt": "",
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 103,
      "reasoning_excerpt": "The user wants the weather in Oakland. I have a get_weather tool. I should call it with city \"Oakland\".",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "request_index": 2,
      "seconds": 5.061,
      "test": "tools",
      "usage": {
        "completion_tokens": 38,
        "prompt_tokens": 4967,
        "total_tokens": 5005
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content": "",
      "content_chars": 0,
      "content_excerpt": "",
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 121,
      "reasoning_excerpt": "The user wants to know the weather in Oakland. I have a get_weather tool available. I should call it with city \"Oakland\".",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "request_index": 1,
      "seconds": 5.13,
      "test": "tools",
      "usage": {
        "completion_tokens": 41,
        "prompt_tokens": 4967,
        "total_tokens": 5008
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content_chars": 0,
      "content_excerpt": "",
      "event_count": 32,
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 122,
      "reasoning_excerpt": "The user wants to know the weather in Oakland. I should use the get_weather tool with the city parameter set to \"Oakland\".",
      "reasoning_field": "stream_delta",
      "reasoning_tokens": null,
      "seconds": 0.651,
      "sse_done": true,
      "test": "streaming-tools",
      "usage": {
        "completion_tokens": 40,
        "prompt_tokens": 166,
        "total_tokens": 206
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content": "",
      "content_chars": 0,
      "content_excerpt": "",
      "finish_reason": "tool_calls",
      "passed": true,
      "reasoning_chars": 123,
      "reasoning_excerpt": "The user wants to know the weather in Oakland. I have a get_weather tool available. Let me call it with the city \"Oakland\".",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "seconds": 0.679,
      "test": "tool-result-initial",
      "usage": {
        "completion_tokens": 42,
        "prompt_tokens": 166,
        "total_tokens": 208
      },
      "validation_detail": "valid tool_call get_weather(city='Oakland')"
    },
    {
      "content": "OAKLAND 72F",
      "content_chars": 11,
      "content_excerpt": "OAKLAND 72F",
      "finish_reason": "stop",
      "passed": true,
      "reasoning_chars": 208,
      "reasoning_excerpt": "The user wants me to reply exactly \"OAKLAND 72F\" using the tool result. The tool result showed Oakland at 72\u00b0F, sunny. They want an exact reply: \"OAKLAND 72F\". I should just output that exactly, nothi",
      "reasoning_field": "reasoning_content",
      "reasoning_tokens": null,
      "seconds": 0.953,
      "test": "tool-result-continuation",
      "usage": {
        "completion_tokens": 64,
        "prompt_tokens": 70,
        "total_tokens": 134
      },
      "validation_detail": "tool result retained"
    }
  ],
  "ocr_expect": [],
  "passed": true,
  "results": [
    {
      "detail": "4.1s got=\"```python\\nsum(xs)\\n```\\n\\nThat's it \u2014 Python's built-\" finish='stop' visible=278 reasoning_chars=711 reasoning_tokens=None",
      "name": "smoke (short coding)",
      "passed": true
    },
    {
      "detail": "parsed keys=['language', 'ok'] finish='stop' visible=31 reasoning_chars=187 reasoning_tokens=None",
      "name": "structured JSON",
      "passed": true
    },
    {
      "detail": "3.4s ctx~8192 got='ZEBRA-42917-QUARTZ' finish='stop' visible=18 reasoning_chars=559 reasoning_tokens=None",
      "name": "needle @ ~8192 ctx",
      "passed": true
    },
    {
      "detail": "4/4 clean (sample: 4.9s valid tool_call get_weather(city='Oakland') finish='tool_calls' visible=0 reasoning_chars=78 reasoning_tokens=None)",
      "name": "shared-prefix tool batch x4",
      "passed": true
    },
    {
      "detail": "0.7s done=True events=32 valid tool_call get_weather(city='Oakland') finish='tool_calls'",
      "name": "streaming tool call",
      "passed": true
    },
    {
      "detail": "1.6s tool result retained finish='stop' visible=11 reasoning_chars=208 reasoning_tokens=None",
      "name": "tool-result continuation",
      "passed": true
    }
  ],
  "schema_version": "preflight/v2",
  "thinking": {
    "chat_template_kwargs": null,
    "mode": "default",
    "reasoning_effort": "max"
  },
  "video_expect": [],
  "video_input": null
}
