Serena
:rocket: Serena — это мощный набор инструментов для создания агента кодирования , способный превратить LLM в полнофункциональный агент, работающий непосредственно с вашей кодовой базой .
:wrench: Serena предоставляет необходимые инструменты для извлечения и редактирования семантического кода , которые аналогичны возможностям IDE, извлекая сущности кода на уровне символов и используя реляционную структуру.
:free: Serena — бесплатная программа с открытым исходным кодом , расширяющая возможности LLM, к которым у вас уже есть бесплатный доступ.
Демонстрация
Вот демонстрация того, как Serena реализует небольшую функцию для себя (лучший GUI журнала) с помощью Claude Desktop. Обратите внимание, как инструменты Serena позволяют Claude находить и редактировать нужные символы.
https://github.com/user-attachments/assets/6eaa9aa1-610d-4723-a2d6-bf1e487ba753
Интеграция LLM
Serena предоставляет необходимые инструменты для кодирования рабочих процессов, но для выполнения фактической работы, организации использования инструментов, требуется степень магистра права.
Интеграция Serena с LLM возможна несколькими способами:
с использованием протокола контекста модели (MCP) .
Serena предоставляет сервер MCP, который интегрируется сКлод Десктоп,
IDE, такие как VSCode, Cursor или IntelliJ,
Расширения, такие как Cline или Roo Code
Goose (для приятного использования CLI)
и многие другие, включая приложение ChatGPT, которое скоро появится
с использованием Agno — фреймворка агента, не зависящего от модели .
Агент Serena на основе Agno позволяет превратить практически любой LLM в агента кодирования, независимо от того, предоставляется ли он Google, OpenAI или Anthropic (с платным ключом API) или бесплатной моделью, предоставляемой Ollama, Together или Anyscale.путем включения инструментов Serena в выбранную вами среду агента.
Реализация инструмента Serena отделена от кода, специфичного для фреймворка, и поэтому может быть легко адаптирована к любому фреймворку агента.
Поддержка языков программирования и возможности семантического анализа
Возможности семантического анализа кода Serena основаны на языковых серверах , использующих широко распространенный протокол языкового сервера (LSP). LSP предоставляет набор универсальных функций запроса и редактирования кода, основанных на символическом понимании кода. Оснащенная этими возможностями, Serena обнаруживает и редактирует код так же, как это делает опытный разработчик, использующий возможности IDE. Serena может эффективно находить нужный контекст и делать правильные вещи даже в очень больших и сложных проектах! Так что он не только бесплатный и с открытым исходным кодом, но и часто достигает лучших результатов, чем существующие решения, которые взимают плату.
Языковые серверы обеспечивают поддержку широкого спектра языков программирования. С Serena мы обеспечиваем
прямая, готовая поддержка для:
Питон
TypeScript/JavaScript
ФП
Go (сначала необходимо установить go и gopls)
Ржавчина
С/С++
Java ( Примечание : запуск медленный, особенно первоначальный. Похоже, что на Mac есть проблемы с Java)
косвенная поддержка (может потребоваться изменение кода/ручная установка) для:
Рубин (непроверено)
C# (непроверено)
Kotlin (непроверено)
Дарт (непроверенный)
Эти языки поддерживаются языковой серверной библиотекой multilspy , которую Serena использует под капотом. Но мы не проверяли явно, работает ли поддержка этих языков на самом деле.
В принципе, можно легко обеспечить поддержку дополнительных языков, предоставив неглубокий адаптер для реализации нового языкового сервера.
Оглавление
Related MCP server: Nabu + Nisaba
Для чего я могу использовать Serena?
Вы можете использовать Serena для любых задач кодирования – будь то анализ, планирование, проектирование новых компонентов или рефакторинг существующих. Поскольку инструменты Serena позволяют LLM замыкать когнитивный цикл восприятия-действия, агенты на основе Serena могут автономно выполнять задачи кодирования от начала до конца – от начального анализа до внедрения, тестирования и, наконец, фиксации системы контроля версий.
Serena может читать, писать и выполнять код, читать логи и вывод терминала. Хотя мы не обязательно поощряем это, "vibe-кодирование" определенно возможно, и если вы хотите почти чувствовать, что "код больше не существует", вы можете найти Serena даже более подходящей для vibing, чем агент внутри IDE (поскольку у вас будет отдельный GUI, который действительно позволит вам забыть).
Бесплатные агенты кодирования с Serena
Даже бесплатный уровень Claude от Anthropic поддерживает MCP-серверы, так что вы можете использовать Serena с Claude бесплатно. Предположительно, то же самое вскоре станет возможным и с ChatGPT Desktop, как только будет добавлена поддержка MCP-серверов.
С помощью Agno у вас также есть возможность использовать Serena с моделью свободных/открытых весов.
Serena — это вклад Oraios AI в сообщество разработчиков.
Мы сами пользуемся им регулярно.
Мы устали от необходимости платить за несколько подписок на основе IDE (например, Windsurf или Cursor), которые заставляли нас продолжать покупать токены сверх уже существующих расходов на подписку на чат. Значительные расходы на API, которые несут такие инструменты, как Claude Code, Cline, Aider и другие инструменты на основе API, также непривлекательны. Поэтому мы создали Serena с перспективой отмены большинства других подписок.
Быстрый старт
Serena можно использовать различными способами, ниже вы найдете инструкции по выборочным интеграциям.
Если вы просто хотите превратить Claude в бесплатный агент кодирования, мы рекомендуем использовать Serena через Claude Desktop.
Если вы хотите использовать Gemini или любую другую модель и вам нужен графический интерфейс, вам следует использовать Agno . На macOS вы также можете использовать графический интерфейс goose .
Если вы предпочитаете использовать Serena через CLI, вы можете использовать goose . Здесь снова возможна почти любая модель.
Если вы хотите использовать Serena, интегрированную в вашу IDE, см. раздел о других клиентах MCP .
Настройка и конфигурирование
Установить
uv(инструкции здесь )Клонируйте репозиторий в
/path/to/serena.Скопируйте
serena_config.template.ymlвserena_config.ymlи настройте параметры.cp serena_config.template.yml serena_config.ymlСкопируйте
project.template.ymlвproject.ymlи настройте параметры, специфичные для вашего проекта (добавьте один такой файл для каждого проекта, над которым вы хотите, чтобы работала Serena). Мы рекомендуем вам скопировать его в каталог.serenaвашего проекта, например,mkdir -p /myproject/.serena cp project.template.yml /myproject/.serena/project.ymlЕсли вы хотите, чтобы Serena динамически переключалась между проектами, добавьте список всех файлов проектов, созданных на предыдущем шаге, в список
projectsвserena_config.yml.
⚠️ Примечание: Serena находится в стадии активной разработки. Мы постоянно добавляем функции, улучшаем стабильность и UX. В результате конфигурация может измениться некорректно. Если у вас недопустимая конфигурация, сервер MCP или агент на базе Serena могут не запуститься (в первом случае изучите журналы MCP). Проверьте журнал изменений и шаблоны конфигурации при обновлении Serena, соответствующим образом адаптируя свои конфигурации.
После первоначальной настройки перейдите к одному из разделов ниже в зависимости от того, как вы хотите использовать Serena.
MCP-сервер (Claude Desktop)
Создайте файл конфигурации для вашего проекта, например
myproject.yml, на основе шаблона в myproject.template.yml .Настройте сервер MCP в вашем клиенте.
Для Claude Desktop (доступно для Windows и macOS) перейдите в Файл / Настройки / Разработчик / Серверы MCP / Изменить конфигурацию, что позволит вам открыть файл JSONclaude_desktop_config.json. Добавьте следующее (с измененными путями), чтобы включить Serena:{ "mcpServers": { "serena": { "command": "/abs/path/to/uv", "args": ["run", "--directory", "/abs/path/to/serena", "serena-mcp-server", "--project", "/abs/path/to/project"] } } }:info: Передача файла проекта необязательна, если вы установили
enable_project_activationв своей конфигурации, поскольку этот параметр позволит вам просто указать Клоду активировать проект, над которым вы хотите работать.Для Claude Desktop (как показано выше) контекст Serena по умолчанию (
desktop-app) и режимы (например,interactive,editing) обычно подходят для общего использования. Обычно вам не нужно указывать их явно вargsесли вы не хотите переопределить значения по умолчанию.Для интеграций IDE (например, VSCode, Cursor, Cline и т. д., настроенных путем добавления Serena в качестве сервера MCP) следует явно передать контекст
ide-assistant, добавив"--context", "ide-assistant"кargsв конфигурации вашего клиента MCP. При желании можно также указать начальные режимы (например,"--mode", "editing").Для определенных одноразовых задач на любом клиенте (например, создание отчета или плана за одно взаимодействие) вы можете указать Serena (после подключения) переключаться в такие режимы, как
planningиone-shot, с помощью инструментаswitch_modesили изначально задать их с помощью флагов--mode, если вы настраиваете команду запуска сервера непосредственно для таких задач.
Более подробную информацию о доступных параметрах и настройках см. в разделе «Режимы и контексты» .
Если вы используете пути, содержащие обратные косые черты, для путей в Windows (обратите внимание, что вы также можете использовать просто прямые косые черты), обязательно правильно экранируйте их (
\\).
Вот и все! Сохраните конфигурацию и перезапустите Claude Desktop.
Поиск неисправностей
Сообщалось, что некоторые конфигурации клиента/ОС/установки вызывают проблемы при использовании Serena со стандартным протоколом stdio , где сервер MCP запускается клиентским приложением. Если у вас возникли такие проблемы, вы можете запустить Serena в режиме sse , запустив, например,
uv run --directory /path/to/serena serena-mcp-server --transport sse --port 9121 --project /path/to/project(опция --project необязательна). Затем настройте клиент для подключения к http://localhost:9121 .
Примечание: для Windows и macOS существуют официальные приложения Claude Desktop от Anthropic, для Linux существует версия с открытым исходным кодом от сообщества .
⚠️ Обязательно полностью закройте приложение Claude Desktop, так как закрытие Claude просто свернёт его в системный трей — по крайней мере, в Windows.
После перезапуска вы должны увидеть инструменты Серены в интерфейсе чата (обратите внимание на маленький значок молотка).
⚠️ Имена инструментов: Claude Desktop (и большинство клиентов MCP) не разрешают имя сервера. Поэтому вам не следует говорить что-то вроде «используйте инструменты Serena». Вместо этого вы можете указать LLM использовать символические инструменты или использовать определенный инструмент, указав его имя. Более того, если вы используете несколько серверов MCP, вы можете получить конфликты имен инструментов , которые приведут к неопределенному поведению. Например, Serena в настоящее время несовместима с сервером Filesystem MCP из-за конфликтов имен инструментов.
ℹ️ Обратите внимание, что серверы MCP, которые используют stdio в качестве протокола, несколько необычны с точки зрения архитектуры клиент/сервер, поскольку сервер обязательно должен быть запущен клиентом, чтобы связь осуществлялась через стандартный поток ввода/вывода сервера. Другими словами, вам не нужно запускать сервер самостоятельно. Клиентское приложение (например, Claude Desktop) заботится об этом и поэтому его необходимо настроить с помощью команды запуска.
Более подробную информацию о серверах MCP с Claude Desktop можно найти в официальном кратком руководстве по началу работы .
Клод Код
Serena — отличный способ сделать Claude Code и дешевле, и мощнее! Мы собираем несколько примеров для этого и уже получили очень положительные отзывы. Пользователи Claude Code могут добавить serena с
claude mcp add serena -- /path/to/uv "run" --directory /path/to/serena serena-mcp-server --project-file /path/to/projectДругие клиенты MCP — Cline, Roo-Code, Cursor, Windsurf и т. д.
Будучи сервером MCP, Serena может быть включена в любой клиент MCP. Та же конфигурация, что и выше, возможно, с небольшими изменениями, специфичными для клиента, должна работать. Большинство популярных существующих помощников кодирования (расширения IDE или IDE, подобные VSCode) принимают подключение к серверам MCP. Для этих интеграций **рекомендуется использовать контекст ide-assistant **, добавляя "--context", "ide-assistant" к args в конфигурации вашего клиента MCP. Включение Serena обычно повышает их производительность, предоставляя им инструменты для символических операций.
В этом случае выставление счетов за использование продолжает контролироваться клиентом по вашему выбору (в отличие от клиента Claude Desktop). Но вы все равно можете захотеть использовать Serena с помощью такого подхода, например, по одной из следующих причин:
Вы уже используете помощника по кодированию (например, Cline или Cursor) и просто хотите сделать его более мощным.
Вы используете Linux и не хотите использовать созданный сообществом Claude Desktop
Вы хотите более тесной интеграции Serena в вашу IDE и не против заплатить за это
Здесь применимы те же соображения, что и при использовании Serena для Claude Desktop (в частности, конфликты имен инструментов).
При использовании в IDE или расширении, которое имеет встроенные взаимодействия AI для кодирования (а это, по сути, все из них), полный набор инструментов Serena может привести к нежелательным взаимодействиям с внутренними инструментами клиента, которые вы как пользователь можете не контролировать. Это особенно касается инструментов редактирования, которые вы можете захотеть отключить для этой цели. По мере накопления опыта использования Serena в различных популярных клиентах мы будем собирать и улучшать лучшие практики, которые обеспечат плавный опыт.
Гусь
goose — это автономный агент кодирования, который имеет интеграцию для серверов MCP и предлагает CLI (и GUI на macOS). Использование goose в настоящее время является самым простым способом запуска Serena через CLI с LLM по вашему выбору.
Для установки следуйте инструкциям здесь .
После этого используйте goose configure для добавления расширения. Для добавления Serena выберите опцию Command-line Extension , назовите его Serena и добавьте следующее в качестве команды:
/abs/path/to/uv run --directory /abs/path/to/serena serena-mcp-server --project /optional/abs/path/to/projectПоскольку Serena может выполнять все необходимые операции редактирования и управления, вам следует отключить расширение developer , которое goose включает по умолчанию. Для этого выполните
goose configureснова выберите опцию Toggle Extensions и убедитесь, что опция Serena включена, а опция developer — нет.
Вот и все. Ознакомьтесь с параметрами конфигурации goose, чтобы узнать, что вы можете с ним сделать (а их много, например, установка различных уровней разрешений для выполнения инструмента).
Goose, похоже, не всегда правильно завершает процессы python для серверов MCP при завершении сеанса. Возможно, вам захочется отключить Serena GUI и/или вручную очистить все запущенные процессы python после завершения работы с goose.
Агно Агент
Agno — это независимая от модели структура агента, которая позволяет превратить Serena в агента (независимо от технологии MCP) с большим количеством базовых LLM. Agno в настоящее время является самым простым способом запуска Serena в графическом интерфейсе чата с LLM по вашему выбору (если вы не используете Mac, то вам может подойти goose, который почти не требует настройки).
Хотя Agno пока не совсем стабилен, мы выбрали его, потому что он поставляется с собственным открытым исходным кодом пользовательского интерфейса, что позволяет легко использовать агента напрямую с помощью интерфейса чата. С Agno Serena превращается в агента (то есть больше не является MCP-сервером), поэтому его можно использовать программными способами (например, для бенчмаркинга или в вашем приложении).
Вот как это работает (см. также документацию Agno ):
Загрузите код агента-ui с помощью npx
npx create-agent-ui@latestили, в качестве альтернативы, клонируйте его вручную:
git clone https://github.com/agno-agi/agent-ui.git cd agent-ui pnpm install pnpm devУстановите serena с дополнительными требованиями:
# You can also only select agno,google or agno,anthropic instead of all-extras uv pip install --all-extras -r pyproject.toml -e .Скопируйте
.env.exampleв.envи заполните ключи API для провайдера(ов), которых вы собираетесь использовать.Запустите приложение agno agent с помощью
uv run python scripts/agno_agent.pyПо умолчанию скрипт использует в качестве модели Клода, но вы можете выбрать любую модель, поддерживаемую Agno (по сути, любую существующую модель).
В новом терминале запустите пользовательский интерфейс agno с помощью
cd agent-ui pnpm devПодключите UI к агенту, который вы запустили выше, и начните общаться. У вас будут те же инструменты, что и в версии сервера MCP.
Ниже представлена короткая демонстрация того, как Серена выполняет небольшую аналитическую задачу с помощью новейшей модели Gemini:
https://github.com/user-attachments/assets/ccfcb968-277d-4ca9-af7f-b84578858c62
⚠️ ВАЖНО: В отличие от подхода MCP-сервера, выполнение инструмента в пользовательском интерфейсе Agno не запрашивает разрешения пользователя. Инструмент оболочки особенно важен, так как он может выполнять произвольное выполнение кода. Хотя мы никогда не сталкивались с какими-либо проблемами с этим в нашем тестировании с Клодом, разрешение этого может быть не совсем безопасным. Вы можете отключить определенные инструменты для своей настройки в файле конфигурации вашего проекта Serena ( .yml ).
Другие агентские фреймворки
Агент Agno особенно хорош благодаря пользовательскому интерфейсу Agno, но Serena легко интегрируется в любую среду агентов (например, pydantic-ai , langgraph или другие).
Вам просто нужно написать адаптер инструментов Serena к инструментам в выбранном вами фреймворке, как это было сделано нами для agno в SerenaAgnoToolkit .
Инструменты и конфигурация Серены
Serena объединяет инструменты для семантического извлечения кода с возможностями редактирования и выполнения оболочки. Поведение Serena можно дополнительно настроить с помощью режимов и контекстов . Полный список инструментов можно найти ниже .
Как правило, рекомендуется использовать все инструменты, поскольку это позволяет Serena обеспечить максимальную отдачу: только выполняя команды оболочки (в частности, тесты), Serena может самостоятельно выявлять и исправлять ошибки.
Однако следует отметить, что инструмент execute_shell_command допускает выполнение произвольного кода. При использовании Serena в качестве сервера MCP клиенты обычно запрашивают у пользователя разрешение перед выполнением инструмента, поэтому, если пользователь заранее проверяет параметры выполнения, это не должно быть проблемой. Однако, если у вас есть опасения, вы можете отключить определенные команды в файле конфигурации .yml вашего проекта. Если вы хотите использовать Serena только для анализа кода и предложения реализаций без изменения кодовой базы, вы можете включить режим только для чтения, установив read_only: true в файле конфигурации вашего проекта. Это автоматически отключит все инструменты редактирования и предотвратит любые изменения в вашей кодовой базе, при этом все возможности анализа и исследования останутся доступными.
В общем, обязательно делайте резервные копии своей работы и используйте систему контроля версий, чтобы избежать потери какой-либо работы.
Сравнение с другими кодирующими агентами
Насколько нам известно, Serena — первый полнофункциональный агент кодирования, все функции которого доступны через сервер MCP, что не требует API-ключей или подписок.
Агенты кодирования на основе подписки
Наиболее известные агенты кодирования на основе подписки являются частями IDE, таких как Windsurf, Cursor и VSCode. Функциональность Serena похожа на Cursor's Agent, Windsurf's Cascade или будущий режим агента VSCode.
Преимущество Serena в том, что она не требует подписки. Потенциальный недостаток в том, что она не интегрирована напрямую в IDE, поэтому проверка нового написанного кода не так проста.
Дополнительные технические различия:
Serena не привязана к конкретной IDE. MCP-сервер Serena может использоваться с любым MCP-клиентом (включая некоторые IDE), а агент на основе Agno предоставляет дополнительные способы применения его функциональности.
Serena не привязана к какой-либо конкретной крупной языковой модели или API.
Serena перемещается и редактирует код с помощью языкового сервера, поэтому имеет символическое понимание кода. Инструменты на основе IDE часто используют подход на основе RAG или чисто текстовый, который часто менее эффективен, особенно для больших кодовых баз.
Serena имеет открытый исходный код и небольшую кодовую базу, поэтому ее можно легко расширять и модифицировать.
Агенты кодирования на основе API
Альтернативой агентам на основе подписки являются агенты на основе API, такие как Claude Code, Cline, Aider, Roo Code и другие, где стоимость использования напрямую соответствует стоимости API базового LLM. Некоторые из них (например, Cline) могут быть даже включены в IDE в качестве расширения. Они часто очень мощные, и их главный недостаток — это (потенциально очень высокая) стоимость API.
Serena сама по себе может использоваться как агент на основе API (см. раздел об Agno выше). Мы пока не написали CLI-инструмент или специальное расширение IDE для Serena (и, вероятно, в последнем нет необходимости, поскольку Serena уже может использоваться с любой IDE, поддерживающей серверы MCP). Если возникнет спрос на Serena как CLI-инструмент, такой как Claude Code, мы рассмотрим возможность его написания.
Главное отличие Serena от других агентов на основе API заключается в том, что Serena может также использоваться как сервер MCP, не требуя при этом ключа API и обходя затраты на API. Это уникальная функция Serena.
Другие кодирующие агенты на основе MCP
Существуют и другие серверы MCP, предназначенные для кодирования, например DesktopCommander и codemcp . Однако, насколько нам известно, ни один из них не предоставляет инструменты для семантического поиска и редактирования кода; они полагаются исключительно на текстовый анализ. Именно интеграция языковых серверов и MCP делает Serena уникальной и такой мощной для сложных задач кодирования, особенно в контексте больших кодовых баз.
Адаптация и воспоминания
По умолчанию Serena выполнит процесс онбординга, когда она впервые запускается для проекта. Целью процесса является ознакомление Serena с проектом и сохранение воспоминаний, которые она затем сможет использовать в будущих взаимодействиях.
Воспоминания — это файлы, хранящиеся в .serena/memories/ в каталоге проекта, которые агент может выбрать для чтения. Можете свободно читать и корректировать их по мере необходимости; вы также можете добавлять новые вручную. Каждый файл в каталоге .serena/memories/ — это файл памяти.
Мы обнаружили, что воспоминания значительно улучшают пользовательский опыт с Сереной. Сама по себе Серена проинструктирована создавать новые воспоминания, когда это уместно.
Режимы и контексты
Поведение и набор инструментов Serena можно настроить с помощью контекстов и режимов . Они обеспечивают высокую степень настройки для наилучшего соответствия вашему рабочему процессу и среде, в которой работает Serena.
Контексты
Контекст определяет общую среду, в которой работает Serena. Он влияет на начальный системный запрос и набор доступных инструментов. Контекст задается при запуске Serena (например, через параметры CLI для сервера MCP или в скрипте агента) и не может быть изменен во время активного сеанса.
Serena поставляется с предопределенными контекстами:
desktop-app: Разработано для использования с приложениями для настольных компьютеров, такими как Claude Desktop. Часто используется по умолчанию.agent: разработан для сценариев, в которых Серена действует как более автономный агент, например, при использовании с Agno.ide-assistant: оптимизирован для интеграции в такие среды разработки, как VSCode, Cursor или Cline, с упором на помощь при кодировании в редакторе.
Вам следует выбрать контекст, который лучше всего соответствует вашей интеграции.
Режимы
Режимы еще больше улучшают поведение Серены для определенных типов задач или стилей взаимодействия. Несколько режимов могут быть активны одновременно, что позволяет вам комбинировать их эффекты. Режимы влияют на системную подсказку и также могут изменять набор доступных инструментов, исключая некоторые из них.
Примеры встроенных режимов включают в себя:
planning: Сосредоточивает Серену на задачах планирования и анализа.editing: оптимизирует Serena для задач прямого изменения кода.interactive: подходит для разговорного стиля взаимодействия.one-shot: настраивает Serena для задач, которые должны быть выполнены за один раз, часто используется приplanningдля создания отчетов или первоначальных планов.no-onboarding: пропускает первоначальный процесс регистрации, если он не нужен для конкретного сеанса.onboarding: (обычно запускается автоматически) фокусируется на процессе онбординга проекта.
Режимы можно устанавливать при запуске (аналогично контекстам), но их также можно переключать динамически во время сеанса. Вы можете поручить LLM использовать инструмент switch_modes для активации другого набора режимов (например, «переключиться на режимы планирования и одноразового использования»).
:warning: Совместимость режимов : хотя вы можете комбинировать режимы, некоторые из них могут быть семантически несовместимыми (например, interactive и one-shot ). В настоящее время Serena не предотвращает несовместимые комбинации; выбор разумных конфигураций режимов остается за пользователем.
Настройка контекстов и режимов
Вы можете создавать собственные контексты и режимы, чтобы точно адаптировать Serena к вашим потребностям:
Добавление в ваш клон Serena : создайте новые файлы
.ymlв каталогахconfig/contexts/илиconfig/modes/в вашем локальном репозитории Serena. Эти пользовательские контексты/режимы будут автоматически зарегистрированы и доступны для использования по имени файла (без расширения.yml). Они также будут отображаться в списках доступных контекстов/режимов.Использование внешних файлов YAML : при запуске Serena вы можете указать абсолютный путь к пользовательскому файлу
.ymlдля контекста или режима.
Файл YAML контекста или режима обычно определяет:
name: (Необязательно, если используется имя файла) Имя контекста/режима.prompt: Строка, которая будет включена в системную подсказку Серены.description: (Необязательно) Краткое описание.excluded_tools: Список названий инструментов (строк), которые следует отключить, когда активен этот контекст/режим.
Такая настройка обеспечивает глубокую интеграцию и адаптацию Serena к конкретным требованиям проекта или личным предпочтениям.
Сочетание с другими серверами MCP
При использовании Serena через MCP Client вы можете использовать его вместе с другими MCP-серверами. Однако остерегайтесь конфликтов имен инструментов! См. информацию об этом выше.
В настоящее время существует конфликт с популярным Filesystem MCP Server. Поскольку Serena также обеспечивает операции с файловой системой, скорее всего, нет необходимости включать эти два одновременно.
Рекомендации по использованию Serena
Мы продолжим собирать лучшие практики по мере роста сообщества Serena. Ниже приведен краткий обзор того, что мы узнали при использовании Serena внутри компании.
Большинство этих рекомендаций справедливы для любого кодирующего агента, включая все агенты, упомянутые выше.
Какую модель выбрать?
К нашему удивлению, Serena, похоже, лучше всего работала с не думающей версией Claude 3.7 по сравнению с думающей версией (мы еще не проводили обширных сравнений с Gemini). Думающая версия работала дольше, испытывала больше трудностей в использовании инструментов и часто просто писала код, не читая достаточно контекста.
В наших первых экспериментах Gemini, похоже, работал очень хорошо. К сожалению, Gemini не поддерживает MCP (пока?), поэтому единственный способ использовать его — через API-ключ. С другой стороны, Gemini сравнительно дешев и может обрабатывать огромные длины контекста.
Адаптация
В самом первом взаимодействии Серене поручено выполнить онбординг и записать первые файлы памяти. Иногда (в зависимости от LLM) файлы не записываются на диск. В этом случае просто попросите Серену записать воспоминания.
На этом этапе Серена обычно читает и пишет довольно много текста и тем самым заполняет контекст. Мы рекомендуем вам переключиться на другой разговор после выполнения онбординга, чтобы не исчерпать токены. Онбординг будет выполнен только один раз, если вы явно не инициируете его.
После регистрации мы рекомендуем вам быстро просмотреть воспоминания и, при необходимости, отредактировать их или добавить дополнительные.
Перед редактированием кода
Лучше всего начинать задачу генерации кода с чистого состояния git. Это не только облегчит вам проверку изменений, но и сама модель получит возможность увидеть, что она изменила, вызвав git diff и тем самым исправить себя или продолжить работу в последующем разговоре, если это необходимо.
:warning: Важно : поскольку Serena будет писать в файлы, используя системные окончания строк, и может захотеть посмотреть на git diff, важно установить git config core.autocrlf в true в Windows. Если git config core.autocrlf установлен в false в Windows, вы можете получить огромные diff только из-за окончаний строк. Обычно рекомендуется включить эту настройку git в Windows:
git config --global core.autocrlf trueВозможные проблемы при редактировании кода
По нашему опыту, LLM действительно плохи в подсчетах, т.е. у них возникают проблемы со вставкой блоков кода в нужное место. Большинство операций редактирования можно выполнить на символическом уровне, что позволяет обойти эту проблему. Однако иногда вставки на уровне строк полезны.
Serena должна дважды проверить номера строк и любые блоки кода, которые она будет редактировать, но вы можете явно указать ей, как редактировать код, если у вас возникнут проблемы.
Исчезновение контекста
Для длительных и сложных задач или задач, где Серена прочитала много контента, вы можете приблизиться к пределам маркеров контекста. В этом случае часто бывает хорошей идеей продолжить в новом разговоре. У Серены есть специальный инструмент для создания сводки текущего состояния прогресса и всей соответствующей информации для его продолжения. Вы можете запросить создание этой сводки и записать ее в память. Затем, в новом разговоре, вы можете просто попросить Серену прочитать память и продолжить задачу. По нашему опыту, это сработало очень хорошо. С другой стороны, поскольку в одном сеансе не задействовано обобщение, Серена обычно не теряется (в отличие от некоторых других агентов, которые обобщают под капотом), и ей также поручено время от времени проверять, находится ли она на правильном пути.
Более того, Серена проинструктирована быть бережливой с контекстом (например, не читать тела символов кода без необходимости), но мы обнаружили, что Клод не всегда очень хорош в бережливости (Gemini, похоже, справляется с этим лучше). Вы можете явно проинструктировать его не читать тела, если знаете, что это не нужно.
Контроль выполнения инструмента
Claude Desktop спросит вас перед запуском инструмента. Для большинства инструментов вы можете просто нажать «Разрешить для этого чата», особенно если все ваши файлы находятся под контролем версий. Исключением является инструмент execute_shell_command — в нем вы можете захотеть проверить каждый вызов по отдельности. Мы рекомендуем просматривать каждый вызов этой команды и не включать ее для всего чата.
Структурирование вашей кодовой базы
Serena использует структуру кода для поиска, чтения и редактирования кода. Это означает, что она будет хорошо работать с хорошо структурированным кодом, но может потерпеть неудачу с полностью неструктурированным (например, класс Бога с огромными немодульными функциями). Аннотации типов также очень помогают здесь. Чем лучше ваш код, тем лучше будет работать Serena. Поэтому мы обычно рекомендуем вам писать хорошо структурированный, модульный и типизированный код — это поможет не только вам, но и вашему ИИ ;).
Ведение журнала, линтинг и тестирование
Serena не может отлаживать (насколько нам известно, на данный момент ни один помощник по кодированию не может этого сделать). Это означает, что для улучшения результатов в цикле агента Serena необходимо получать информацию, выполняя тесты, скрипты, выполняя линтинг и т. д. Часто бывает очень полезно включать много сообщений журнала с явной информацией и иметь содержательные тесты. Особенно последние часто помогают агенту самокорректироваться.
Обычно мы рекомендуем начинать задачу редактирования из состояния, когда все проверки и тесты линтинга пройдены.
Общие советы
Мы обнаружили, что часто бывает полезно потратить некоторое время на концептуализацию и планирование задачи перед ее фактической реализацией, особенно для нетривиальной задачи. Это помогает как в достижении лучших результатов, так и в повышении чувства контроля и пребывания в курсе. Вы можете составить подробный план за один сеанс, где Серена может прочитать большую часть вашего кода, чтобы создать контекст, а затем продолжить реализацию в другом (возможно, после создания подходящих воспоминаний).
Поиск неисправностей
Поддержка серверов MCP в Claude Desktop и различных пакетах SDK для серверов MCP — это относительно новые разработки, которые могут работать нестабильно.
Рабочая конфигурация сервера MCP может различаться от платформы к платформе и от клиента к клиенту. Мы рекомендуем всегда использовать абсолютные пути, так как относительные пути могут быть источниками ошибок. Языковой сервер работает в отдельном подпроцессе и вызывается с помощью asyncio – иногда клиент может привести к его сбою. Если у вас включено окно журнала Serena, и оно исчезает, вы поймете, что произошло.
Некоторые клиенты (например, goose) могут некорректно завершать работу серверов MCP. Следите за зависшими процессами Python и завершайте их вручную, если это необходимо.
Серена Логгинг
Чтобы помочь с устранением неполадок, мы написали небольшую утилиту GUI для ведения журнала. Для большинства клиентов мы рекомендуем включить ее через конфигурацию проекта ( project.yml ), если у вас возникнут проблемы. Многие клиенты также пишут журналы MCP, которые могут помочь в выявлении проблем.
GUI для ведения журнала может работать не для всех клиентов и не во всех системах. В настоящее время он не работает на macOS или в расширениях VSCode, таких как Cline.
Благодарности
Мы создали Serena на основе множества существующих технологий с открытым исходным кодом, наиболее важными из которых являются:
multilspy . Прекрасно спроектированная оболочка вокруг языковых серверов, следующая LSP. Ее было нелегко расширить с помощью символической логики, которая требовалась Serena, поэтому вместо того, чтобы включить ее как зависимость, мы скопировали исходный код и адаптировали его под наши нужды.
Agno и связанный с ним агент-ui , который мы используем, чтобы позволить Serena работать с любой моделью, помимо тех, которые поддерживают MCP.
Все языковые серверы, которые мы используем через multilspy.
Без этих проектов строительство Серены было бы невозможным (или его было бы значительно сложнее построить).
Настройка Серены
Очень легко расширить функциональность искусственного интеллекта Serena с помощью собственных идей. Просто реализуйте новый инструмент, создав подкласс serena.agent.Tool и реализуйте метод apply (не часть интерфейса, см. комментарий в Tool ). По умолчанию SerenaAgent сразу получит к нему доступ.
Также относительно просто добавить поддержку нового языка . Мы с нетерпением ждем, что предложит сообщество! Подробности о том, как внести свой вклад, см. здесь .
Полный список инструментов
Вот полный список инструментов Serena с кратким описанием (вывод uv run serena-list-tools ):
activate_project: Активирует проект по имени.check_onboarding_performed: проверяет, была ли уже выполнена адаптация проекта.create_text_file: создает/перезаписывает файл в каталоге проекта.delete_lines: Удаляет диапазон строк в файле.delete_memory: Удаляет воспоминание из хранилища памяти проекта Серены.execute_shell_command: выполняет команду оболочки.find_referencing_code_snippets: Находит фрагменты кода, в которых упоминается символ в указанном месте.find_referencing_symbols: находит символы, ссылающиеся на символ в указанном месте (возможно фильтрование по типу).find_symbol: выполняет глобальный (или локальный) поиск символов с/содержащими заданное имя/подстроку (возможно фильтрование по типу).get_active_project: получает имя текущего активного проекта (если есть) и выводит список существующих проектовget_current_config: выводит текущую конфигурацию агента, включая активные режимы, инструменты и контекст.get_symbols_overview: Получает обзор символов верхнего уровня, определенных в указанном файле или каталоге.initial_instructions: Получает начальные инструкции для текущего проекта. Следует использовать только в настройках, где системная подсказка не может быть установлена, например, в клиентах, над которыми вы не имеете контроля, таких как Claude Desktop.insert_after_symbol: вставляет содержимое после окончания определения заданного символа.insert_at_line: вставляет содержимое в указанную строку файла.insert_before_symbol: вставляет содержимое перед началом определения заданного символа.list_dir: Выводит список файлов и каталогов в указанном каталоге (возможно с рекурсией).list_memories: Перечисляет воспоминания в хранилище памяти проекта Серены.onboarding: выполняет адаптацию (определяет структуру проекта и основные задачи, например, для тестирования или сборки).prepare_for_new_conversation: предоставляет инструкции по подготовке к новому разговору (чтобы продолжить с необходимым контекстом).read_file: читает файл в каталоге проекта.read_memory: считывает память с указанным именем из хранилища памяти проекта Serena.replace_lines: заменяет диапазон строк в файле новым содержимым.replace_symbol_body: Заменяет полное определение символа.restart_language_server: Перезапускает языковой сервер, может потребоваться при внесении изменений не через Serena.search_for_pattern: Выполняет поиск шаблона в проекте.summarize_changes: Содержит инструкции по обобщению изменений, внесенных в кодовую базу.switch_modes: активирует режимы, предоставляя список их названийthink_about_collected_information: Инструмент для размышления о полноте собранной информации.think_about_task_adherence: инструмент для размышлений, позволяющий определить, находится ли агент на пути к выполнению текущей задачи.think_about_whether_you_are_done: Инструмент для размышлений, позволяющий определить, действительно ли выполнена задача.write_memory: записывает именованную память (для дальнейшего использования) в хранилище памяти проекта Serena.
Available Tools
29 toolsactivate_projectActivate ProjectBRead-only
Activates the project with the given name or path.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | The name of a registered project to activate or a path to a project directory. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, which the description does not contradict. The description adds minimal behavioral context by implying activation of a project, but it doesn't elaborate on effects like environment changes or permissions needed. With annotations covering safety, the description provides some value but not rich behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded and wastes no space, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (activation operation), high schema coverage, annotations, and the presence of an output schema, the description is reasonably complete. It covers the basic action but could benefit from more context on outcomes or integration with sibling tools, though the structured data reduces the burden on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the 'project' parameter. The description adds no additional meaning beyond the schema, such as examples or constraints, but since the schema is comprehensive, a baseline score of 3 is appropriate as the description doesn't compensate unnecessarily.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('activates') and the resource ('the project with the given name or path'), making the purpose specific and understandable. However, it does not differentiate this tool from sibling tools like 'switch_modes' or 'get_current_config', which might relate to project state changes, so it doesn't fully distinguish from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as 'switch_modes' or 'get_current_config', nor does it mention prerequisites like needing a registered project. It lacks explicit when/when-not instructions or named alternatives, leaving usage context unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_onboarding_performedCheck Onboarding PerformedARead-only
Checks whether project onboarding was already performed. You should always call this tool before beginning to actually work on the project/after activating a project.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, which the description aligns with by implying a non-destructive check. The description adds value by specifying the tool's role in workflow sequencing (before work/after activation), but doesn't provide additional behavioral details like error handling or output interpretation beyond what annotations cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose followed by usage guidelines. Every word serves a clear function, with no redundancy or unnecessary elaboration, making it highly efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, annotations covering safety (read-only, non-destructive), and an output schema (implied by context signals), the description is reasonably complete. It explains what the tool does and when to use it, though it could benefit from hinting at the output's meaning (e.g., boolean result or status details) to fully compensate for lack of output schema explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so no parameter documentation is needed. The description doesn't mention parameters, which is appropriate. A baseline of 4 is applied since no parameters exist, and the description focuses correctly on the tool's purpose and usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Checks whether project onboarding was already performed.' It specifies the verb ('checks') and resource ('project onboarding'), making the intent unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'onboarding' or 'activate_project', which might have overlapping contexts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: 'You should always call this tool before beginning to actually work on the project/after activating a project.' This gives clear timing and context for when to use it, including a reference to the sibling tool 'activate_project' as a related action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_text_fileCreate Text FileADestructive
Write a new file or overwrite an existing file. Returns a message indicating success or failure.
| Name | Required | Description | Default |
|---|---|---|---|
| relative_path | Yes | The relative path to the file to create. | |
| content | Yes | The (appropriately encoded) content to write to the file. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds valuable context by specifying that it can 'overwrite an existing file' and returns 'success or failure' messages. This clarifies the destructive nature beyond the annotation and provides outcome expectations, though it doesn't mention permissions, rate limits, or file encoding details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two tightly focused sentences that efficiently convey the core functionality and outcome. Every word serves a purpose with zero redundancy, and the information is front-loaded with the primary action stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive nature indicated by annotations, the presence of an output schema, and 100% parameter coverage, the description provides adequate context. It covers the tool's primary behavior and outcome expectations, though it could benefit from mentioning encoding requirements or error scenarios for a more complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the input schema fully documents both parameters. The description doesn't add any parameter-specific information beyond what's in the schema, so it meets the baseline score of 3. No additional semantic context is provided for 'relative_path' or 'content' parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Write a new file or overwrite an existing file') and resource ('file'), distinguishing it from sibling tools like 'read_file' or 'replace_content'. It precisely communicates both creation and overwrite capabilities in a single concise statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for file creation/overwriting but provides no explicit guidance on when to use this tool versus alternatives like 'replace_content' or 'write_memory'. It mentions the tool's behavior but doesn't specify scenarios where it's preferred over other file manipulation tools in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_memoryDelete MemoryADestructive
Delete a memory file. Should only happen if a user asks for it explicitly, for example by saying that the information retrieved from a memory file is no longer correct or no longer relevant for the project.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_file_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, which cover the core safety profile. The description adds valuable context beyond this by emphasizing the need for explicit user consent and providing rationale examples, which helps the agent understand the gravity of the operation. It doesn't detail error conditions or output behavior, but with annotations present, this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences that are front-loaded with the core action and immediately followed by critical usage guidelines. Every word serves a purpose, with no redundancy or fluff, making it highly efficient and easy to parse for an AI agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's destructive nature, the description effectively complements the annotations by stressing user consent and providing examples. With an output schema present, it doesn't need to explain return values. The combination of clear purpose, strong guidelines, and annotations makes this description sufficiently complete for safe agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, with only one parameter ('memory_file_name') documented structurally. The description doesn't add any parameter-specific information, such as format examples or constraints. However, with a single parameter and the tool's clear purpose, the baseline of 3 is appropriate as the schema alone provides minimal but adequate coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Delete') and resource ('a memory file'), making the purpose immediately understandable. It distinguishes from sibling tools like 'edit_memory', 'read_memory', and 'write_memory' by specifying deletion rather than modification or retrieval. However, it doesn't explicitly contrast with 'list_memories' or other file-related tools beyond the verb itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: 'only if a user asks for it explicitly', with concrete examples ('information retrieved from a memory file is no longer correct or no longer relevant for the project'). This clearly defines the appropriate context and helps the agent avoid misuse, especially given the destructive nature implied by the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_memoryEdit MemoryADestructive
Replaces content matching a regular expression in a memory.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_file_name | Yes | The name of the memory. | |
| needle | Yes | The string or regex pattern to search for. If `mode` is "literal", this string will be matched exactly. If `mode` is "regex", this string will be treated as a regular expression (syntax of Python's `re` module, with flags DOTALL and MULTILINE enabled). | |
| repl | Yes | The replacement string (verbatim). | |
| mode | Yes | Either "literal" or "regex", specifying how the `needle` parameter is to be interpreted. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, indicating this is a mutation tool with destructive potential. The description adds useful context about regex matching behavior (Python re module with DOTALL/MULTILINE flags) and the replacement being verbatim, which goes beyond annotations. However, it doesn't mention error conditions, side effects, or what happens when no match is found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core functionality ('Replaces content matching a regular expression in a memory'). Every word earns its place with no redundancy or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (regex editing with destructive potential), the description provides adequate context when combined with rich annotations and a complete input schema. However, it could benefit from mentioning the existence of an output schema (which handles return values) and providing more behavioral context about edge cases. The combination of description, annotations, and schema is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, all parameters are well-documented in the schema itself. The description doesn't add significant semantic information beyond what's already in the parameter descriptions, which thoroughly explain memory_file_name, needle (with mode-specific behavior), repl, and mode. The baseline of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Replaces content') on a specific resource ('in a memory') using a specific method ('matching a regular expression'). It distinguishes from siblings like 'replace_content' (general file replacement) and 'write_memory' (full overwrite) by specifying regex-based partial editing of memory files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for regex-based editing of memory content, but doesn't explicitly state when to use this vs alternatives like 'replace_content' for non-memory files or 'write_memory' for complete overwrites. It provides clear context about the operation type but lacks explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_shell_commandExecute Shell CommandADestructive
Execute a shell command and return its output. If there is a memory about suggested commands, read that first. Never execute unsafe shell commands! IMPORTANT: Do not use this tool to start
long-running processes (e.g. servers) that are not intended to terminate quickly,
processes that require user interaction. Returns a JSON object containing the command's stdout and optionally stderr output.
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | The shell command to execute. | |
| cwd | No | The working directory to execute the command in. If None, the project root will be used. | |
| capture_stderr | No | Whether to capture and return stderr output. | |
| max_answer_chars | No | If the output is longer than this number of characters, no content will be returned. -1 means using the default value, don't adjust unless there is no other way to get the content required for the task. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds valuable behavioral context beyond this: it warns against unsafe commands, specifies output truncation behavior via max_answer_chars, mentions checking memory first, and describes the JSON return structure. This provides important safety and operational guidance that annotations alone don't cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the core purpose, followed by important warnings and return format details. While some sentences could be more concise (e.g., the warning about long-running processes is slightly verbose), overall it's efficient with each sentence serving a clear purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (destructive shell command execution), rich annotations (destructiveHint=true), complete schema coverage, and existence of an output schema, the description provides excellent contextual completeness. It covers safety warnings, usage constraints, memory integration, and output behavior, making it fully adequate for an AI agent to understand when and how to use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the input schema already fully documents all 4 parameters. The description doesn't add any parameter-specific semantics beyond what's in the schema descriptions, so it meets the baseline expectation without providing additional value about parameter usage or interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('execute a shell command and return its output') and distinguishes it from siblings by focusing on command execution rather than file operations, memory management, or project configuration. It goes beyond just restating the name/title by specifying the return format and behavioral constraints.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when NOT to use this tool (for long-running processes or processes requiring user interaction) and references checking memory for suggested commands first. However, it doesn't explicitly name alternative tools for those excluded use cases or differentiate from similar tools like list_dir for directory operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_fileFind FileARead-only
Finds non-gitignored files matching the given file mask within the given relative path. Returns a JSON object with the list of matching files.
| Name | Required | Description | Default |
|---|---|---|---|
| file_mask | Yes | The filename or file mask (using the wildcards * or ?) to search for. | |
| relative_path | Yes | The relative path to the directory to search in; pass "." to scan the project root. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds useful behavioral context by specifying 'non-gitignored files' (exclusion behavior) and the return format ('JSON object with the list of matching files'), but does not mention potential limitations like recursion depth or performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core functionality and includes essential details (exclusion of gitignored files, return format) without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations (readOnlyHint, destructiveHint), 100% schema coverage, and presence of an output schema, the description is complete enough. It covers purpose, key behavioral trait (non-gitignored), and return format, leaving detailed parameter and output documentation to the structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters well-documented in the schema. The description adds minimal value beyond the schema by implying the search scope and exclusion of gitignored files, but does not provide additional syntax or format details for the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Finds'), resource ('non-gitignored files'), and scope ('matching the given file mask within the given relative path'), distinguishing it from siblings like 'list_dir' (which lists directory contents) and 'search_for_pattern' (which searches file content).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (searching for files by name/mask, excluding gitignored files), but does not explicitly state when not to use it or name alternatives like 'list_dir' for directory listing or 'search_for_pattern' for content search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_referencing_symbolsFind Referencing SymbolsARead-only
Finds references to the symbol at the given name_path. The result will contain metadata about the referencing symbols
as well as a short code snippet around the reference. Returns a list of JSON objects with the symbols referencing the requested symbol.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path | Yes | For finding the symbol to find references for, same logic as in the `find_symbol` tool. | |
| relative_path | Yes | The relative path to the file containing the symbol for which to find references. Note that here you can't pass a directory but must pass a file. | |
| include_kinds | No | Same as in the `find_symbol` tool. | |
| exclude_kinds | No | Same as in the `find_symbol` tool. | |
| max_answer_chars | No | Same as in the `find_symbol` tool. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds useful behavioral context beyond annotations by specifying what the result contains (metadata about referencing symbols and short code snippets) and that it returns a list of JSON objects. However, it doesn't mention potential limitations like performance impacts or result size constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences with zero waste. The first sentence states the core functionality and result format, while the second clarifies the return type. Every word contributes to understanding the tool's purpose and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of annotations (readOnlyHint, destructiveHint), 100% schema coverage, and an output schema (implied by 'Returns a list of JSON objects'), the description provides complete contextual information. It adequately explains what the tool does, what it returns, and references sibling tools where appropriate, making it sufficient for agent understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the input schema already documents all 5 parameters thoroughly, including references to the 'find_symbol' tool for parameter behavior. The description doesn't add significant semantic information beyond what's in the schema, maintaining the baseline score of 3 for adequate but not enhanced parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Finds references') and resource ('the symbol at the given name_path'), and distinguishes it from sibling tools by specifying it returns referencing symbols rather than finding symbols themselves. It explicitly mentions what the result contains (metadata and code snippets), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by mentioning the tool finds references to a symbol, suggesting it should be used when you need to know where a symbol is referenced. However, it doesn't explicitly state when to use this tool versus alternatives like 'find_symbol' or provide exclusion criteria, though the parameter descriptions reference 'find_symbol' for some parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_symbolFind SymbolARead-only
Retrieves information on all symbols/code entities (classes, methods, etc.) based on the given name path pattern.
The returned symbol information can be used for edits or further queries.
Specify depth > 0 to also retrieve children/descendants (e.g., methods of a class).
A name path is a path in the symbol tree within a source file.
For example, the method my_method defined in class MyClass would have the name path MyClass/my_method.
If a symbol is overloaded (e.g., in Java), a 0-based index is appended (e.g. "MyClass/my_method[0]") to
uniquely identify it.
To search for a symbol, you provide a name path pattern that is used to match against name paths. It can be
a simple name (e.g. "method"), which will match any symbol with that name
a relative path like "class/method", which will match any symbol with that name path suffix
an absolute name path "/class/method" (absolute name path), which requires an exact match of the full name path within the source file. Append an index
[i]to match a specific overload only, e.g. "MyClass/my_method[1]". Returns a list of symbols (with locations) matching the name.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path_pattern | Yes | The name path matching pattern (see above). | |
| depth | No | Depth up to which descendants shall be retrieved (e.g. use 1 to also retrieve immediate children; for the case where the symbol is a class, this will return its methods). Default 0. | |
| relative_path | No | Optional. Restrict search to this file or directory. If None, searches entire codebase. If a directory is passed, the search will be restricted to the files in that directory. If a file is passed, the search will be restricted to that file. If you have some knowledge about the codebase, you should use this parameter, as it will significantly speed up the search as well as reduce the number of results. | |
| include_body | No | If True, include the symbol's source code. Use judiciously. | |
| include_kinds | No | Optional. List of LSP symbol kind integers to include. (e.g., 5 for Class, 12 for Function). Valid kinds: 1=file, 2=module, 3=namespace, 4=package, 5=class, 6=method, 7=property, 8=field, 9=constructor, 10=enum, 11=interface, 12=function, 13=variable, 14=constant, 15=string, 16=number, 17=boolean, 18=array, 19=object, 20=key, 21=null, 22=enum member, 23=struct, 24=event, 25=operator, 26=type parameter. If not provided, all kinds are included. | |
| exclude_kinds | No | Optional. List of LSP symbol kind integers to exclude. Takes precedence over `include_kinds`. If not provided, no kinds are excluded. | |
| substring_matching | No | If True, use substring matching for the last element of the pattern, such that "Foo/get" would match "Foo/getValue" and "Foo/getData". | |
| max_answer_chars | No | Max characters for the JSON result. If exceeded, no content is returned. -1 means the default value from the config will be used. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds valuable behavioral context beyond annotations: it explains that results can be used for edits, describes the matching logic (simple name, relative path, absolute path), mentions overload handling with indices, and notes performance considerations (using relative_path speeds up search). It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with the core purpose, followed by detailed explanations of name paths and patterns. While comprehensive, it is appropriately sized for a complex tool with many parameters. Some sentences could be more concise (e.g., the name path explanation is lengthy), but overall it avoids unnecessary repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, 100% schema coverage, annotations, and an output schema), the description is complete. It covers purpose, usage examples, behavioral nuances (like overload handling and performance tips), and does not need to explain return values since an output schema exists. It adequately complements the structured data without gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds some semantic context for name_path_pattern (explaining patterns and examples) and depth (linking it to retrieving children), but most parameter details are already in the schema. This meets the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('retrieves information on all symbols/code entities') and distinguishes it from siblings by focusing on symbol lookup rather than file operations (find_file), pattern searching (search_for_pattern), or symbol editing (rename_symbol, replace_symbol_body). It explicitly mentions what the returned information can be used for ('for edits or further queries'), which helps differentiate its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (e.g., 'Specify `depth > 0` to also retrieve children/descendants') and implies alternatives through sibling tool names like find_file or search_for_pattern, but it does not explicitly state when not to use it or name specific alternatives. The guidance on using the relative_path parameter for speed and reduced results offers practical usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_current_configGet Current ConfigARead-only
Print the current configuration of the agent, including the active and available projects, tools, contexts, and modes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, covering safety. The description adds valuable context by specifying what configuration components are included (projects, tools, contexts, modes), which isn't inferable from annotations alone. However, it doesn't mention output format details or potential limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action ('Print the current configuration') and then enumerates included components. Every word adds value with zero redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, read-only, non-destructive), the description fully covers its purpose and scope. With annotations providing safety context and an output schema existing (so return values needn't be described), the description is complete for this low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters and 100% schema description coverage, the baseline is 4. The description appropriately doesn't discuss parameters since none exist, avoiding unnecessary detail while clearly indicating the tool operates without inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Print') and the exact resource ('current configuration of the agent'), listing all included components (active/available projects, tools, contexts, modes). It distinguishes itself from siblings like 'get_symbols_overview' or 'list_memories' by focusing on overall agent configuration rather than specific data subsets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing to inspect the agent's configuration state, but provides no explicit guidance on when to use this versus alternatives like 'check_onboarding_performed' or 'switch_modes'. It doesn't specify prerequisites, exclusions, or comparative contexts with sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_symbols_overviewGet Symbols OverviewARead-only
Use this tool to get a high-level understanding of the code symbols in a file. This should be the first tool to call when you want to understand a new file, unless you already know what you are looking for. Returns a JSON object containing info about top-level symbols in the file.
| Name | Required | Description | Default |
|---|---|---|---|
| relative_path | Yes | The relative path to the file to get the overview of. | |
| depth | No | Depth up to which descendants of top-level symbols shall be retrieved (e.g. 1 retrieves immediate children). Default 0. | |
| max_answer_chars | No | If the overview is longer than this number of characters, no content will be returned. -1 means the default value from the config will be used. Don't adjust unless there is really no other way to get the content required for the task. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds valuable context beyond this by specifying that it returns 'a JSON object containing info about top-level symbols in the file,' which clarifies the output format. However, it doesn't mention potential limitations like rate limits or error conditions, leaving some behavioral aspects uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in three sentences: the first states the purpose, the second provides usage guidelines, and the third describes the return format. Each sentence adds essential value without redundancy, making it front-loaded and appropriately concise for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations (readOnlyHint, destructiveHint), 100% schema coverage, and the presence of an output schema, the description is complete enough. It covers purpose, usage context, and output format, leaving detailed parameter and return value documentation to the structured fields, which is appropriate for this setup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all three parameters. The description adds no additional parameter semantics beyond what's in the schema (e.g., it doesn't explain 'relative_path' context or 'depth' implications further). This meets the baseline of 3 since the schema carries the full burden, but the description doesn't compensate with extra insights.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('get a high-level understanding of the code symbols in a file') and resource ('file'), distinguishing it from sibling tools like 'find_symbol' or 'read_file' by focusing on symbol overview rather than searching or raw content. It explicitly positions this as the first tool for understanding new files, making its purpose distinct and well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('first tool to call when you want to understand a new file') and when not to ('unless you already know what you are looking for'), effectively differentiating it from alternatives like 'find_symbol' for targeted searches. This clear contextual advice helps the agent choose appropriately among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
initial_instructionsInitial InstructionsARead-only
Provides the 'Serena Instructions Manual', which contains essential information on how to use the Serena toolbox. IMPORTANT: If you have not yet read the manual, call this tool immediately after you are given your task by the user, as it will critically inform you!.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds that it provides 'essential information' and is 'critical' to inform the agent, which adds context about importance. However, it doesn't specify format, length, or structure of the manual content beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with two sentences that each serve distinct purposes: stating what the tool provides and giving critical usage instructions. It's front-loaded with the core purpose, though the second sentence could be slightly more streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, annotations covering safety, and an output schema exists, the description provides adequate context about purpose and critical timing. It doesn't need to explain return values since output schema exists, but could optionally hint at manual content format for slightly better completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters and 100% schema description coverage, the baseline is 4. The description appropriately doesn't discuss parameters since none exist, focusing instead on the tool's purpose and usage timing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'provides the Serena Instructions Manual' which is a specific resource, though it doesn't specify the exact content or format of the manual. It distinguishes itself from siblings by being the only tool focused on providing essential instructions rather than performing project operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: 'If you have not yet read the manual, call this tool immediately after you are given your task by the user, as it will critically inform you!' This clearly indicates when to use it (immediately after task assignment if manual not read) and implies it's foundational before using other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
insert_after_symbolInsert After SymbolADestructive
Inserts the given body/content after the end of the definition of the given symbol (via the symbol's location). A typical use case is to insert a new class, function, method, field or variable assignment.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path | Yes | Name path of the symbol after which to insert content (definitions in the `find_symbol` tool apply). | |
| relative_path | Yes | The relative path to the file containing the symbol. | |
| body | Yes | The body/content to be inserted. The inserted code shall begin with the next line after the symbol. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds valuable behavioral context: it specifies that insertion occurs 'after the end of the definition' and that 'the inserted code shall begin with the next line after the symbol.' This clarifies the exact positioning behavior beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences: the first states the core functionality, the second provides a typical use case. Every word earns its place, and the most important information (what the tool does) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of both annotations (destructiveHint=true, readOnlyHint=false) and an output schema (implied by context signals), the description provides complete contextual information. It covers the tool's purpose, typical usage, and behavioral specifics without needing to explain return values or safety characteristics that are already documented elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, all parameters are already documented in the input schema. The description adds some context by mentioning 'symbol's location' and referencing 'find_symbol' for name_path, but doesn't provide significant additional semantic meaning beyond what's in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('inserts') and target ('after the end of the definition of the given symbol'), with explicit mention of the content being inserted ('body/content'). It distinguishes from sibling 'insert_before_symbol' by specifying 'after' positioning, and from 'replace_symbol_body' by indicating insertion rather than replacement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('to insert a new class, function, method, field or variable assignment') and references the 'find_symbol' tool for determining symbol locations. However, it doesn't explicitly state when NOT to use it or directly compare with alternatives like 'insert_before_symbol' or 'replace_content'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
insert_before_symbolInsert Before SymbolADestructive
Inserts the given content before the beginning of the definition of the given symbol (via the symbol's location). A typical use case is to insert a new class, function, method, field or variable assignment; or a new import statement before the first symbol in the file.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path | Yes | Name path of the symbol before which to insert content (definitions in the `find_symbol` tool apply). | |
| relative_path | Yes | The relative path to the file containing the symbol. | |
| body | Yes | The body/content to be inserted before the line in which the referenced symbol is defined. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true and readOnlyHint=false, which the description aligns with by describing an insertion operation that modifies files. The description adds valuable context beyond annotations by specifying that insertion occurs 'before the beginning of the definition' and via 'the symbol's location', and mentions typical use cases, which helps the agent understand the tool's behavior in practical scenarios.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by a second sentence providing typical use cases. Both sentences earn their place by clarifying scope and practical applications without redundancy or unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (file modification with symbol-based positioning), the description provides sufficient context alongside annotations (destructive, not read-only) and a complete input schema. With an output schema present, the description does not need to explain return values, making it complete for agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing clear documentation for all three parameters (name_path, relative_path, body). The description adds minimal semantic value beyond the schema, only implying that 'body' is content to insert and referencing 'find_symbol' for name_path definitions. This meets the baseline of 3 when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('inserts') and resource ('content before the beginning of the definition of the given symbol'), with specific examples of typical use cases (new class, function, method, field, variable assignment, or import statement). It distinguishes from sibling 'insert_after_symbol' by specifying 'before' rather than 'after'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('to insert a new class, function, method, field or variable assignment; or a new import statement before the first symbol in the file'), but does not explicitly state when not to use it or name alternatives beyond the implied sibling 'insert_after_symbol'. It lacks explicit exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dirList DirARead-only
Lists files and directories in the given directory (optionally with recursion). Returns a JSON object with the names of directories and files within the given directory.
| Name | Required | Description | Default |
|---|---|---|---|
| relative_path | Yes | The relative path to the directory to list; pass "." to scan the project root. | |
| recursive | Yes | Whether to scan subdirectories recursively. | |
| skip_ignored_files | No | Whether to skip files and directories that are ignored. | |
| max_answer_chars | No | If the output is longer than this number of characters, no content will be returned. -1 means the default value from the config will be used. Don't adjust unless there is really no other way to get the content required for the task. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds value by specifying the return format ('JSON object with names of directories and files') and hinting at recursion behavior, but does not disclose additional traits like rate limits, auth needs, or error conditions beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence and adds useful details in the second. Both sentences earn their place by clarifying functionality and output format without redundancy or unnecessary elaboration, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations (readOnlyHint, destructiveHint), and the presence of an output schema, the description is largely complete. It covers purpose, optional recursion, and return format, though it could benefit from more explicit usage guidelines or edge-case handling to be fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds minimal semantics by mentioning recursion and the return format, but does not provide extra details on parameter usage or interactions beyond what the schema already covers, aligning with the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Lists') and resource ('files and directories'), specifies the scope ('in the given directory'), and mentions an optional feature ('with recursion'). It distinguishes itself from sibling tools like 'find_file' or 'search_for_pattern' by focusing on directory listing rather than searching or filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for directory listing but does not explicitly state when to use this tool versus alternatives like 'find_file' or 'search_for_pattern'. It mentions recursion as an option but lacks guidance on scenarios where recursion is preferred or when to avoid it, leaving usage context somewhat implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_memoriesList MemoriesARead-only
List available memories. Any memory can be read using the read_memory tool.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds no behavioral traits beyond this, such as pagination, sorting, or access constraints, relying entirely on annotations for safety disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with two sentences that are front-loaded and waste-free. Every word contributes to understanding the tool's purpose and usage, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, read-only, non-destructive) and the presence of annotations and an output schema, the description is complete enough for basic use. It could benefit from more detail on output format or limitations, but the essentials are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters and 100% schema description coverage, the schema fully documents the lack of inputs. The description doesn't need to add parameter details, and its mention of 'available memories' implies no filtering, which aligns with the empty schema. Baseline is 4 for zero parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('List') and resource ('memories'), making it immediately understandable. However, it doesn't explicitly differentiate from sibling tools like 'read_memory' beyond mentioning it as a follow-up action, missing direct comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by stating that listed memories can be read with 'read_memory', implying usage as a precursor to that tool. It doesn't specify when not to use it or alternatives, but the guidance is sufficient for basic navigation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onboardingOnboardingARead-only
Call this tool if onboarding was not performed yet. You will call this tool at most once per conversation. Returns instructions on how to create the onboarding information.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds context about the one-time-per-conversation constraint and that it returns instructions, which are useful behavioral details beyond the annotations. However, it doesn't describe error handling or what happens if called multiple times.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and well-structured with two sentences: the first states when to call it, and the second specifies the call frequency and return value. Every sentence adds essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, annotations covering safety, and an output schema (implied by context signals), the description is mostly complete. It covers purpose, usage guidelines, and behavioral constraints. However, it could briefly mention what the instructions entail or link to sibling tools for more context, but the output schema likely handles return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing instead on usage context. A baseline of 4 is applied since there are no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to be called when onboarding hasn't been performed yet, and it returns instructions for creating onboarding information. It specifies the verb 'call' and the resource 'onboarding', but doesn't explicitly differentiate from sibling tools like 'check_onboarding_performed' or 'initial_instructions' beyond the conditional trigger.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: 'Call this tool if onboarding was not performed yet' and 'You will call this tool at most once per conversation.' This clearly defines when to use it (onboarding not done) and includes a usage constraint (once per conversation), though it doesn't name specific alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_for_new_conversationPrepare For New ConversationBRead-only
Instructions for preparing for a new conversation. This tool should only be called on explicit user request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe, non-destructive operation. The description adds no behavioral context beyond what annotations provide, such as what 'preparing' entails or any side effects. However, it doesn't contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with zero wasted words. It's appropriately sized and front-loaded, though the first sentence is uninformative. Every sentence serves a purpose: the first states the tool's name, and the second provides critical usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has zero parameters, annotations covering safety, and an output schema (which means return values are documented elsewhere), the description is minimally adequate. However, it fails to explain what 'preparing for a new conversation' actually means or what the tool does, leaving a significant gap in understanding its purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and schema description coverage is 100%. With no parameters to document, the description doesn't need to compensate for any gaps. The baseline for zero parameters is 4, as there's nothing to explain beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Instructions for preparing for a new conversation' is a tautology that restates the tool's name/title without specifying what the tool actually does. It lacks a clear verb+resource combination and doesn't distinguish this tool from its many siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states 'This tool should only be called on explicit user request,' providing clear when-to-use guidance. This is a strong, unambiguous usage rule that helps the agent avoid inappropriate invocations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileRead FileARead-only
Reads the given file or a chunk of it. Generally, symbolic operations like find_symbol or find_referencing_symbols should be preferred if you know which symbols you are looking for. Returns the full text of the file at the given relative path.
| Name | Required | Description | Default |
|---|---|---|---|
| relative_path | Yes | The relative path to the file to read. | |
| start_line | No | The 0-based index of the first line to be retrieved. | |
| end_line | No | The 0-based index of the last line to be retrieved (inclusive). If None, read until the end of the file. | |
| max_answer_chars | No | If the file (chunk) is longer than this number of characters, no content will be returned. Don't adjust unless there is really no other way to get the content required for the task. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds valuable context beyond this: it explains that it can read 'a chunk' of a file (via start_line/end_line parameters) and warns about the max_answer_chars constraint ('no content will be returned' if exceeded). This enhances behavioral understanding without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences: the first states the core functionality and preferred alternatives, the second clarifies the return value. Every sentence serves a clear purpose with zero wasted words, making it easy to parse and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (4 parameters, 1 required), 100% schema coverage, annotations covering safety, and an output schema (implied by 'Returns...'), the description is complete. It covers purpose, guidelines, and key behavioral aspects without needing to repeat schema details or explain return values extensively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all four parameters. The description mentions 'chunk' reading and the max_answer_chars behavior, but these details are already covered in the schema descriptions for start_line, end_line, and max_answer_chars. It adds minimal semantic value beyond what the structured schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Reads') and resource ('the given file or a chunk of it'), distinguishing it from sibling tools like find_symbol or find_referencing_symbols. It explicitly mentions what it returns ('full text of the file at the given relative path'), making the purpose unambiguous and well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use alternatives: 'Generally, symbolic operations like find_symbol or find_referencing_symbols should be preferred if you know which symbols you are looking for.' This clearly indicates when not to use this tool and names specific sibling alternatives, offering strong contextual direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_memoryRead MemoryARead-only
Read the content of a memory file. This tool should only be used if the information is relevant to the current task. You can infer whether the information is relevant from the memory file name. You should not read the same memory file multiple times in the same conversation.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_file_name | Yes | ||
| max_answer_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating this is a safe read operation. The description adds valuable behavioral context beyond annotations: it specifies relevance criteria (based on file name) and a usage constraint (no repeated reads in same conversation). However, it doesn't disclose other potential behaviors like error handling, response format, or performance characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in three sentences: the first states the core purpose, the second provides usage criteria, and the third adds a behavioral constraint. Every sentence adds value without redundancy, and it's front-loaded with the essential action. No wasted words or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (a read operation with relevance filtering), annotations cover safety (read-only, non-destructive), and an output schema exists (so return values needn't be described), the description is reasonably complete. It adds useful context like relevance criteria and usage limits, though it lacks parameter explanations and doesn't fully address sibling tool differentiation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema provides no parameter documentation. The description doesn't explain either parameter's semantics—it mentions 'memory file name' but doesn't clarify its format or source, and omits 'max_answer_chars' entirely. Since parameters are few (2) and one has a default, the baseline is 3, but the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Read') and resource ('content of a memory file'), making the purpose unambiguous. It distinguishes this tool from siblings like 'list_memories' (which lists files) and 'write_memory' (which writes content). However, it doesn't explicitly contrast with 'read_file' (which reads general files), leaving some sibling differentiation incomplete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('if the information is relevant to the current task') and when not to use it ('should not read the same memory file multiple times in the same conversation'). It also implies alternatives by referencing the memory file name for relevance inference, though it doesn't name specific sibling tools like 'list_memories' for discovery.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_symbolRename SymbolADestructive
Renames the symbol with the given name_path to new_name throughout the entire codebase.
Note: for languages with method overloading, like Java, name_path may have to include a method's
signature to uniquely identify a method. Returns result summary indicating success or failure.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path | Yes | Name path of the symbol to rename (definitions in the `find_symbol` tool apply). | |
| relative_path | Yes | The relative path to the file containing the symbol to rename. | |
| new_name | Yes | The new name for the symbol. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds valuable context: it specifies the scope ('throughout the entire codebase'), mentions language-specific considerations (Java method overloading), and describes the return format ('result summary indicating success or failure'). This goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly sized at three sentences, front-loaded with the core purpose, followed by important implementation notes and return value information. Every sentence earns its place with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive nature (annotations), 3 parameters with full schema coverage, and the existence of an output schema, the description provides complete context. It covers purpose, scope, language considerations, and return format without needing to explain parameters or output details already documented elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents all three parameters thoroughly. The description adds minimal extra context: it references 'find_symbol' tool for name_path definitions and mentions method signatures for overloading, but doesn't provide significant additional parameter semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('renames') and resource ('symbol with the given name_path'), specifying scope ('throughout the entire codebase'). It distinguishes from siblings like 'replace_symbol_body' by focusing on renaming rather than content replacement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (renaming symbols across codebases) and includes a note about method overloading in languages like Java. However, it doesn't explicitly state when NOT to use it or name specific alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_contentReplace ContentADestructive
Replaces one or more occurrences of a given pattern in a file with new content.
This is the preferred way to replace content in a file whenever the symbol-level tools are not appropriate.
VERY IMPORTANT: The "regex" mode allows very large sections of code to be replaced without fully quoting them!
Use a regex of the form "beginning.*?end-of-text-to-be-replaced" to be faster and more economical!
ALWAYS try to use wildcards to avoid specifying the exact content to be replaced,
especially if it spans several lines. Note that you cannot make mistakes, because if the regex should match
multiple occurrences while you disabled allow_multiple_occurrences, an error will be returned, and you can retry
with a revised regex.
Therefore, using regex mode with suitable wildcards is usually the best choice!.
| Name | Required | Description | Default |
|---|---|---|---|
| relative_path | Yes | The relative path to the file. | |
| needle | Yes | The string or regex pattern to search for. If `mode` is "literal", this string will be matched exactly. If `mode` is "regex", this string will be treated as a regular expression (syntax of Python's `re` module, with flags DOTALL and MULTILINE enabled). | |
| repl | Yes | The replacement string (verbatim). If mode is "regex", the string can contain backreferences to matched groups in the needle regex, specified using the syntax $!1, $!2, etc. for groups 1, 2, etc. | |
| mode | Yes | Either "literal" or "regex", specifying how the `needle` parameter is to be interpreted. | |
| allow_multiple_occurrences | No | If True, the regex may match multiple occurrences in the file and all of them will be replaced. If this is set to False and the regex matches multiple occurrences, an error will be returned (and you may retry with a revised, more specific regex). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true and readOnlyHint=false, which the description aligns with by describing a replacement operation. The description adds valuable behavioral context beyond annotations: it explains error handling when allow_multiple_occurrences is false, provides regex performance optimization tips, and mentions retry capabilities. However, it doesn't cover rate limits or specific permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately front-loaded with the core purpose, but contains some redundancy in advocating for regex mode (mentioned multiple times with similar points). The 'VERY IMPORTANT' section could be more streamlined, though all sentences contribute meaningful guidance about tool usage strategies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (destructive file operation with regex capabilities), the description provides comprehensive context: it explains when to use this versus alternatives, offers detailed regex usage strategies, describes error behavior, and references sibling tools. With annotations covering safety aspects and an output schema presumably handling return values, the description fills all necessary contextual gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds meaningful context about parameter usage: it emphasizes regex mode advantages for large sections, explains wildcard strategies to avoid exact content specification, and clarifies the interaction between regex patterns and the allow_multiple_occurrences parameter. This provides practical guidance beyond the schema's technical specifications.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('replaces one or more occurrences of a given pattern in a file with new content') and distinguishes it from sibling tools by mentioning 'symbol-level tools' as alternatives. It explicitly names the resource (file content) and operation (replacement).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('preferred way to replace content... whenever the symbol-level tools are not appropriate') and offers detailed advice on regex mode usage versus literal mode. It also references sibling tools like 'replace_symbol_body' as alternatives for symbol-level operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_symbol_bodyReplace Symbol BodyADestructive
Replaces the body of the symbol with the given name_path.
The tool shall be used to replace symbol bodies that have been previously retrieved
(e.g. via find_symbol).
IMPORTANT: Do not use this tool if you do not know what exactly constitutes the body of the symbol.
| Name | Required | Description | Default |
|---|---|---|---|
| name_path | Yes | For finding the symbol to replace, same logic as in the `find_symbol` tool. | |
| relative_path | Yes | The relative path to the file containing the symbol. | |
| body | Yes | The new symbol body. The symbol body is the definition of a symbol in the programming language, including e.g. the signature line for functions. IMPORTANT: The body does NOT include any preceding docstrings/comments or imports, in particular. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true and readOnlyHint=false, which the description aligns with by implying a mutation ('Replaces'). The description adds valuable context beyond annotations: it clarifies that the body excludes 'preceding docstrings/comments or imports,' specifies a prerequisite (previous retrieval via find_symbol), and warns about misuse if the body is unclear. This enhances behavioral understanding without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by usage guidelines and a critical warning. Each sentence earns its place by providing essential information without redundancy, resulting in a well-structured and efficient text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (destructive mutation with 3 required parameters), the description is complete: it covers purpose, usage context, prerequisites, and critical warnings. With annotations providing safety cues and an output schema present (implying return values are documented elsewhere), no additional information is needed for effective tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing detailed parameter documentation (e.g., 'body' includes the definition excluding docstrings). The description adds minimal semantics beyond the schema, such as linking 'name_path' to 'find_symbol' logic, but does not significantly enhance parameter understanding. With high schema coverage, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Replaces the body of the symbol') and identifies the target resource ('symbol with the given name_path'). It distinguishes from siblings like 'rename_symbol' (which changes the name) and 'replace_content' (which replaces file content rather than symbol bodies), establishing a unique purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use ('to replace symbol bodies that have been previously retrieved via find_symbol') and when not to use ('Do not use this tool if you do not know what exactly constitutes the body of the symbol'). It also references a specific alternative tool ('find_symbol') for preparation, offering clear context for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_for_patternSearch For PatternARead-only
Offers a flexible search for arbitrary patterns in the codebase, including the possibility to search in non-code files. Generally, symbolic operations like find_symbol or find_referencing_symbols should be preferred if you know which symbols you are looking for.
Pattern Matching Logic: For each match, the returned result will contain the full lines where the substring pattern is found, as well as optionally some lines before and after it. The pattern will be compiled with DOTALL, meaning that the dot will match all characters including newlines. This also means that it never makes sense to have .* at the beginning or end of the pattern, but it may make sense to have it in the middle for complex patterns. If a pattern matches multiple lines, all those lines will be part of the match. Be careful to not use greedy quantifiers unnecessarily, it is usually better to use non-greedy quantifiers like .*? to avoid matching too much content.
File Selection Logic:
The files in which the search is performed can be restricted very flexibly.
Using restrict_search_to_code_files is useful if you are only interested in code symbols (i.e., those
symbols that can be manipulated with symbolic tools like find_symbol).
You can also restrict the search to a specific file or directory,
and provide glob patterns to include or exclude certain files on top of that.
The globs are matched against relative file paths from the project root (not to the relative_path parameter that
is used to further restrict the search).
Smartly combining the various restrictions allows you to perform very targeted searches. Returns A mapping of file paths to lists of matched consecutive lines.
| Name | Required | Description | Default |
|---|---|---|---|
| substring_pattern | Yes | Regular expression for a substring pattern to search for. | |
| context_lines_before | No | Number of lines of context to include before each match. | |
| context_lines_after | No | Number of lines of context to include after each match. | |
| paths_include_glob | No | Optional glob pattern specifying files to include in the search. Matches against relative file paths from the project root (e.g., "*.py", "src/**/*.ts"). Supports standard glob patterns (*, ?, [seq], **, etc.) and brace expansion {a,b,c}. Only matches files, not directories. If left empty, all non-ignored files will be included. | |
| paths_exclude_glob | No | Optional glob pattern specifying files to exclude from the search. Matches against relative file paths from the project root (e.g., "*test*", "**/*_generated.py"). Supports standard glob patterns (*, ?, [seq], **, etc.) and brace expansion {a,b,c}. Takes precedence over paths_include_glob. Only matches files, not directories. If left empty, no files are excluded. | |
| relative_path | No | Only subpaths of this path (relative to the repo root) will be analyzed. If a path to a single file is passed, only that will be searched. The path must exist, otherwise a `FileNotFoundError` is raised. | |
| restrict_search_to_code_files | No | Whether to restrict the search to only those files where analyzed code symbols can be found. Otherwise, will search all non-ignored files. Set this to True if your search is only meant to discover code that can be manipulated with symbolic tools. For example, for finding classes or methods from a name pattern. Setting to False is a better choice if you also want to search in non-code files, like in html or yaml files, which is why it is the default. | |
| max_answer_chars | No | If the output is longer than this number of characters, no content will be returned. -1 means the default value from the config will be used. Don't adjust unless there is really no other way to get the content required for the task. Instead, if the output is too long, you should make a stricter query. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, which the description aligns with by describing a search operation. The description adds significant behavioral context beyond annotations, detailing pattern matching logic (e.g., DOTALL compilation, line inclusion, greedy vs. non-greedy quantifiers) and file selection logic (e.g., glob patterns, restrictions), though it doesn't explicitly mention rate limits or auth needs, which are not contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections ('Pattern Matching Logic', 'File Selection Logic') and front-loaded key information. It is appropriately sized, but some sentences could be more concise (e.g., the explanation of DOTALL and greedy quantifiers is slightly verbose), though overall it avoids waste and is easy to follow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (8 parameters, regex patterns, file restrictions) and the presence of annotations and an output schema (implied by 'Returns A mapping of file paths to lists of matched consecutive lines'), the description is complete. It covers usage scenarios, behavioral details, and parameter interactions without needing to explain return values, making it sufficient for an AI agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds some value by explaining the purpose of parameters like 'restrict_search_to_code_files' and how glob patterns work relative to the project root, but it doesn't provide significant additional semantics beyond what's in the schema, warranting a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Offers a flexible search for arbitrary patterns in the codebase, including the possibility to search in non-code files.' It specifies the verb ('search'), resource ('patterns in the codebase'), and scope ('including non-code files'), and distinguishes it from sibling tools by explicitly mentioning alternatives like 'find_symbol' and 'find_referencing_symbols'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool versus alternatives: 'Generally, symbolic operations like find_symbol or find_referencing_symbols should be preferred if you know which symbols you are looking for.' It also advises on context, such as using 'restrict_search_to_code_files' for code symbols and setting it to 'False' for non-code files, offering clear alternatives and exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
switch_modesSwitch ModesBRead-only
Activates the desired modes, like ["editing", "interactive"] or ["planning", "one-shot"].
| Name | Required | Description | Default |
|---|---|---|---|
| modes | Yes | The names of the modes to activate. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe, non-destructive operation. The description adds minimal context by implying activation of modes, but doesn't disclose behavioral traits like what 'activation' entails (e.g., state changes, side effects, or interactions with other tools). It doesn't contradict annotations, but offers little beyond them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose with examples. It avoids unnecessary words and gets straight to the point. However, it could be slightly more structured by explicitly stating the tool's role in the context of sibling tools, but as-is, it's concise and well-formed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which handles return values), annotations covering safety, and high schema coverage, the description is minimally adequate. It explains the basic action but lacks context on what modes are, how they interact with other tools, or when to use this. For a tool that likely changes system state (despite readOnlyHint), more detail on behavior and usage would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'modes' fully documented in the schema. The description adds value by providing examples (e.g., ['editing', 'interactive']), which clarify the expected format and possible values beyond the schema's generic array of strings. However, it doesn't explain semantics like what modes are available or their effects, keeping it at the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Activates') and resource ('the desired modes'), making the purpose understandable. It provides specific examples like 'editing', 'interactive', 'planning', and 'one-shot' which help illustrate what modes might be. However, it doesn't explicitly differentiate from sibling tools like 'activate_project' or 'get_current_config', which could cause confusion about when to use this versus those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, context for activation, or exclusions. Given sibling tools like 'activate_project' and 'get_current_config', the lack of differentiation leaves the agent without clear usage rules, relying solely on the tool name and description which are vague about scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
think_about_collected_informationThink About Collected InformationARead-only
Think about the collected information and whether it is sufficient and relevant. This tool should ALWAYS be called after you have completed a non-trivial sequence of searching steps like find_symbol, find_referencing_symbols, search_files_for_pattern, read_file, etc.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe, non-destructive operation. The description adds context about when to call it (after searching steps), which is useful behavioral guidance beyond the annotations. However, it doesn't disclose details like what the tool actually does (e.g., returns analysis, triggers internal processing) or any rate limits, leaving some behavioral aspects unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences that are front-loaded with the core purpose and followed by specific usage guidelines. Every sentence adds value without redundancy, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, annotations cover safety (read-only, non-destructive), an output schema exists (so return values are documented elsewhere), and the description provides clear usage context, it's mostly complete. However, it could be more explicit about what the tool outputs or how it aids decision-making, slightly reducing completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so there's no need for parameter details in the description. The description appropriately doesn't discuss parameters, which is efficient. A baseline of 4 is applied since no parameters exist, and the description doesn't introduce unnecessary complexity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool is for 'thinking about collected information' and assessing sufficiency/relevance, which gives a general purpose. However, it's somewhat vague about what specific action the tool performs (e.g., does it analyze, summarize, or just prompt reflection?) and doesn't clearly differentiate from sibling tools like 'think_about_task_adherence' or 'think_about_whether_you_are_done' beyond the information focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it states the tool 'should ALWAYS be called after you have completed a non-trivial sequence of searching steps' and lists examples like 'find_symbol, find_referencing_symbols, search_files_for_pattern, read_file, etc.' This clearly defines when to use it versus alternatives, making it highly actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
think_about_task_adherenceThink About Task AdherenceARead-only
Think about the task at hand and whether you are still on track. Especially important if the conversation has been going on for a while and there has been a lot of back and forth.
This tool should ALWAYS be called before you insert, replace, or delete code.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, which already convey that this is a safe, non-destructive operation. The description adds context about its role in ensuring task adherence before code modifications, which is useful behavioral insight beyond the annotations. However, it does not detail aspects like potential side effects, performance implications, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, with three sentences that directly address purpose and usage without unnecessary details. Each sentence adds value, such as the specific trigger conditions and mandatory usage before code changes, making it well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (low, with 0 parameters), rich annotations (readOnlyHint, destructiveHint), and the presence of an output schema, the description is complete enough. It explains the tool's role in task adherence and when to use it, which aligns with the structured data. However, it could slightly enhance completeness by mentioning what the output schema returns, though this is not strictly necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so there is no need for parameter explanation in the description. The description appropriately does not discuss parameters, which is efficient, but since there are no parameters, a baseline of 4 is applied as it fully compensates for the lack of parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool's purpose as 'Think about the task at hand and whether you are still on track,' which is a clear action but somewhat vague in scope. It does not specify what resources or data it operates on, and it does not distinguish itself from sibling tools like 'think_about_collected_information' or 'think_about_whether_you_are_done,' making it less specific than ideal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: 'Especially important if the conversation has been going on for a while and there has been a lot of back and forth' and 'This tool should ALWAYS be called before you insert, replace, or delete code.' This gives clear context for when to use it, though it does not explicitly state when not to use it or compare it to alternatives like other 'think' tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
think_about_whether_you_are_doneThink About Whether You Are DoneARead-only
Whenever you feel that you are done with what the user has asked for, it is important to call this tool.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true and destructiveHint=false, which the description does not contradict. The description adds context by implying this is a self-assessment or decision-making tool, but it does not disclose additional behavioral traits like what triggers the 'done' state, any side effects, or how it interacts with other tools. With annotations covering safety, the description provides minimal extra value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose and usage. It is front-loaded with the key information and contains no redundant or extraneous details, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, annotations covering read-only and non-destructive behavior, and an output schema exists, the description is reasonably complete. It explains when to use the tool, which is the core need. However, it could be more detailed about what 'done' entails or how it integrates with other tools, leaving minor gaps in full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description does not add parameter details, which is appropriate. A baseline of 4 is applied since there are no parameters, and the description does not introduce unnecessary complexity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool's purpose as 'call this tool whenever you feel that you are done with what the user has asked for,' which is a clear action but vague in scope. It specifies a verb ('call') but lacks concrete details about what 'done' means or what resource it operates on, making it somewhat ambiguous compared to more specific sibling tools like 'create_text_file' or 'read_file'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use the tool: 'whenever you feel that you are done with what the user has asked for.' This gives a clear context for invocation. However, it does not mention when not to use it or name alternatives, such as sibling tools like 'think_about_collected_information' or 'think_about_task_adherence,' which might be relevant for related decision-making processes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_memoryWrite MemoryADestructive
Write some information (utf-8-encoded) about this project that can be useful for future tasks to a memory in md format. The memory name should be meaningful.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_file_name | Yes | ||
| content | Yes | ||
| max_answer_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true and readOnlyHint=false, which the description aligns with by implying a write operation. The description adds valuable context beyond annotations: it specifies the memory is stored in markdown format and should have a meaningful name, which helps the agent understand behavioral expectations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences that front-load the core purpose and key requirements. Every word adds value: the first sentence covers action, content, format, and purpose; the second emphasizes naming quality. No wasted words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a destructive tool with 3 parameters (0% schema coverage) and an output schema, the description adequately covers the purpose and format but lacks parameter details and behavioral nuances like error handling or idempotency. The output schema may help with return values, but the description doesn't fully compensate for the low parameter coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'memory name should be meaningful' (hinting at memory_file_name) and 'information... to a memory' (hinting at content), but doesn't explain max_answer_chars or provide details on parameter formats, constraints, or interactions. This leaves significant gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Write some information') and resource ('to a memory'), specifying the format ('md format') and encoding ('utf-8-encoded'). It distinguishes from siblings like 'read_memory' and 'edit_memory' by focusing on creation, but doesn't explicitly differentiate from 'create_text_file' which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context ('about this project', 'useful for future tasks'), suggesting when to use it for project documentation. However, it lacks explicit guidance on when to choose this over alternatives like 'create_text_file' or 'edit_memory', and doesn't mention prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
17 tool updates
v1.0.0- Changed
create_text_file1 field changed- changed
Input schema / properties / content / descriptionPrevious value: -"The (utf-8-encoded) content to write to the file."New value: +"The (appropriately encoded) content to write to the file."
- Added
edit_memory - Changed
execute_shell_command2 fields changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - changed
Input schema / properties / max_answer_chars / descriptionPrevious value: -"If the output is longer than this number of characters,\nno content will be returned. Don't adjust unless there is really no other way to get the content\nrequired for the task."New value: +"If the output is longer than this number of characters,\nno content will be returned. -1 means using the default value, don't adjust unless there is no other way to get the content\nrequired for the task."
- Changed
find_referencing_symbols1 field changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1
- Changed
find_symbol7 fields changed- changed
Input schema / properties / depth / descriptionPrevious value: -"Depth to retrieve descendants (e.g., 1 for class methods/attributes)."New value: +"Depth up to which descendants shall be retrieved (e.g. use 1 to also retrieve immediate children;\nfor the case where the symbol is a class, this will return its methods).\nDefault 0." - changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - changed
Input schema / properties / max_answer_chars / descriptionPrevious value: -"Max characters for the JSON result. If exceeded, no content is returned."New value: +"Max characters for the JSON result. If exceeded, no content is returned.\n-1 means the default value from the config will be used." - removed
Input schema / properties / name_pathRemoved value: -{ - "description": "The name path pattern to search for, see above for details.", - "title": "Name Path", - "type": "string" -} - added
Input schema / properties / name_path_patternAdded value: +{ + "description": "The name path matching pattern (see above).", + "title": "Name Path Pattern", + "type": "string" +} - changed
Input schema / properties / substring_matching / descriptionPrevious value: -"If True, use substring matching for the last segment of `name`."New value: +"If True, use substring matching for the last element of the pattern, such that\n\"Foo/get\" would match \"Foo/getValue\" and \"Foo/getData\"." - changed
Input schema / requiredPrevious value: -[ - "name_path" -]New value: +[ + "name_path_pattern" +]
- Added
get_current_config - Changed
get_symbols_overview3 fields changed- added
Input schema / properties / depthAdded value: +{ + "default": 0, + "description": "Depth up to which descendants of top-level symbols shall be retrieved\n(e.g. 1 retrieves immediate children). Default 0.", + "title": "Depth", + "type": "integer" +} - changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - changed
Input schema / properties / max_answer_chars / descriptionPrevious value: -"If the overview is longer than this number of characters,\nno content will be returned. Don't adjust unless there is really no other way to get the content\nrequired for the task."New value: +"If the overview is longer than this number of characters,\nno content will be returned. -1 means the default value from the config will be used.\nDon't adjust unless there is really no other way to get the content required for the task."
- Added
initial_instructions - Changed
list_dir3 fields changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - changed
Input schema / properties / max_answer_chars / descriptionPrevious value: -"If the output is longer than this number of characters,\nno content will be returned. Don't adjust unless there is really no other way to get the content\nrequired for the task."New value: +"If the output is longer than this number of characters,\nno content will be returned. -1 means the default value from the config will be used.\nDon't adjust unless there is really no other way to get the content required for the task." - added
Input schema / properties / skip_ignored_filesAdded value: +{ + "default": false, + "description": "Whether to skip files and directories that are ignored.", + "title": "Skip Ignored Files", + "type": "boolean" +}
- Changed
read_file1 field changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1
- Changed
read_memory1 field changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1
- Added
rename_symbol - Added
replace_content - Removed
replace_regex - Changed
replace_symbol_body1 field changed- changed
Input schema / properties / body / descriptionPrevious value: -"The new symbol body. Important: Begin directly with the symbol definition and provide no\nleading indentation for the first line (but do indent the rest of the body according to the context)."New value: +"The new symbol body. The symbol body is the definition of a symbol\nin the programming language, including e.g. the signature line for functions.\nIMPORTANT: The body does NOT include any preceding docstrings/comments or imports, in particular."
- Changed
search_for_pattern4 fields changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - changed
Input schema / properties / max_answer_chars / descriptionPrevious value: -"If the output is longer than this number of characters,\nno content will be returned. Don't adjust unless there is really no other way to get the content\nrequired for the task. Instead, if the output is too long, you should\nmake a stricter query."New value: +"If the output is longer than this number of characters,\nno content will be returned.\n-1 means the default value from the config will be used.\nDon't adjust unless there is really no other way to get the content\nrequired for the task. Instead, if the output is too long, you should\nmake a stricter query." - changed
Input schema / properties / paths_exclude_glob / descriptionPrevious value: -"Optional glob pattern specifying files to exclude from the search.\nMatches against relative file paths from the project root (e.g., \"*test*\", \"**/*_generated.py\").\nTakes precedence over paths_include_glob. Only matches files, not directories. If left empty, no files are excluded."New value: +"Optional glob pattern specifying files to exclude from the search.\nMatches against relative file paths from the project root (e.g., \"*test*\", \"**/*_generated.py\").\nSupports standard glob patterns (*, ?, [seq], **, etc.) and brace expansion {a,b,c}.\nTakes precedence over paths_include_glob. Only matches files, not directories. If left empty, no files are excluded." - changed
Input schema / properties / paths_include_glob / descriptionPrevious value: -"Optional glob pattern specifying files to include in the search.\nMatches against relative file paths from the project root (e.g., \"*.py\", \"src/**/*.ts\").\nOnly matches files, not directories. If left empty, all non-ignored files will be included."New value: +"Optional glob pattern specifying files to include in the search.\nMatches against relative file paths from the project root (e.g., \"*.py\", \"src/**/*.ts\").\nSupports standard glob patterns (*, ?, [seq], **, etc.) and brace expansion {a,b,c}.\nOnly matches files, not directories. If left empty, all non-ignored files will be included."
- Changed
write_memory4 fields changed- changed
Input schema / properties / max_answer_chars / defaultPrevious value: -200000New value: +-1 - added
Input schema / properties / memory_file_nameAdded value: +{ + "title": "Memory File Name", + "type": "string" +} - removed
Input schema / properties / memory_nameRemoved value: -{ - "title": "Memory Name", - "type": "string" -} - changed
Input schema / requiredPrevious value: -[ - "memory_name", - "content" -]New value: +[ + "memory_file_name", + "content" +]
25 tool updates
- First observed
activate_project - First observed
check_onboarding_performed - First observed
create_text_file - First observed
delete_memory - First observed
execute_shell_command - First observed
find_file - First observed
find_referencing_symbols - First observed
find_symbol - First observed
get_symbols_overview - First observed
insert_after_symbol - First observed
insert_before_symbol - First observed
list_dir - First observed
list_memories - First observed
onboarding - First observed
prepare_for_new_conversation - First observed
read_file - First observed
read_memory - First observed
replace_regex - First observed
replace_symbol_body - First observed
search_for_pattern - First observed
switch_modes - First observed
think_about_collected_information - First observed
think_about_task_adherence - First observed
think_about_whether_you_are_done - First observed
write_memory
TDQS
Most tools have distinct purposes, but there is some overlap between search tools (find_symbol, find_referencing_symbols, search_for_pattern) and file operations (read_file vs. find_file vs. list_dir). The descriptions help clarify differences, but an agent might occasionally misselect between similar search or file access tools.
The naming follows a consistent verb_noun pattern (e.g., activate_project, create_text_file, execute_shell_command) with only minor deviations like initial_instructions (adjective_noun) and think_about_* tools (verb_phrase). Overall, the pattern is predictable and readable.
With 29 tools, the count feels heavy for a code assistant server, bordering on overwhelming. While many tools are specialized (e.g., multiple think_about_* tools), the high number could lead to confusion or inefficiency in tool selection.
The toolset provides comprehensive coverage for code editing and project management, including CRUD operations for files and symbols, search capabilities, memory management, and workflow guidance (e.g., onboarding, thinking tools). No obvious gaps are present for its intended domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Coding agents build full-stack apps in persistent workspaces and share them by link.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseBqualityCmaintenanceA coding agent toolkit that transforms LLMs into coding assistants capable of working directly on your codebase with semantic code retrieval and editing tools, providing IDE-like capabilities without requiring API subscriptions.331MIT
- AlicenseNot gradedqualityBmaintenanceProvides semantic code intelligence tools (search, structural views) and a workspace TUI interface for LLM agents to efficiently navigate codebases, manage context, and maintain architectural patterns across Python, Java, C++, and Perl projects.4MIT
- AlicenseBqualityDmaintenanceA coding agent toolkit that provides IDE-like semantic code retrieval and editing tools, enabling LLMs to efficiently navigate and modify codebases using symbol-level operations instead of basic file reading and string replacements.19MIT
- FlicenseNot gradedqualityDmaintenanceProvides Cursor-like code intelligence using tools like ripgrep, ctags, and tree-sitter to help LLMs explore and understand entire codebases. It implements a structured, phase-gated workflow to ensure high-confidence code modifications and eliminate hallucinations.-
Appeared in Searches
- Free and open-source coding assistants
- Searching for FOSS coding tools without AI hallucinations, supporting all languages, and cross-platform
- FiveM C# Development with AI-Enhanced Codebase Understanding and Reasoning
- A server for finding coding resources and programming help
- Free tools for file management, project structuring, planning, and coding
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/oraios/serena'
If you have feedback or need assistance with the MCP directory API, please join our Discord server