take_screenshot
Capture the current screen via the Dittu phone camera. Returns a JPEG image. Call this to see what's on screen before acting.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Capture the current screen via the Dittu phone camera. Returns a JPEG image. Call this to see what's on screen before acting.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Changes observed during successful MCP inspections. Dates show when Glama detected each change.
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavior. It states it returns a JPEG image but does not explicitly say it is read-only or non-destructive, nor does it discuss permissions or potential failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff, front-loaded with the main action. Every word is useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description mentions the return format (JPEG) but does not explain when to use this tool over closely related siblings. For a simple tool, it is adequate but incomplete in distinguishing from alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and the input schema is fully covered. The description adds no parameter information, which is acceptable for zero-parameter tools. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it captures the current screen via the Dittu phone camera and returns a JPEG image. However, it does not explicitly distinguish this tool from sibling tools like take_screenshot_with_crosshair or take_screenshot_with_grid.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says 'Call this to see what's on screen before acting,' providing a usage context. But it does not specify when not to use it or mention alternative tools (e.g., get_raw_screenshot for raw data).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.
Each tool targets a distinct action: clicks, screenshots (raw, corrected, calibration aids), keyboard input, mouse movement, scrolling, and corner calibration. There is no ambiguity between tools.
Tool names mix conventions: some are single verbs (click, scroll, key), some are verb_noun (move_mouse, set_corners, type_text), and some are longer phrases (take_screenshot_with_crosshair). Although readable, the inconsistency is notable.
With 12 tools, the server covers all essential interactions for camera-based screen control without being bloated or sparse. Each tool has a clear role.
The tool set covers clicking, typing, key combinations, mouse movement, scrolling, and multiple screenshot modes including calibration aids. Missing drag/swipe functionality is a minor gap.