Background / Current Behavior
In watsonx Orchestrate chat, end users can give a thumbs up or thumbs down to an agent response, select a reason, and optionally enter a comment. Currently, the details panel of the conversation view shows the number of positive and negative ratings and the comments for each conversation.
When we specified a conversation, the AgentOps agent estimated the cause of the negative rating from the conversation content and suggested fixes. It answered that it can suggest fixes for Instructions, descriptions of tools and collaborator agents, and knowledge, but can apply changes only to Instructions.
However, on October 6, 2026, we confirmed the following:
When we asked it to summarize thumbs-down comments for a specific agent and explain the causes, it could not identify the relevant conversations.
When we specified the conversation ID, it confirmed that the rating was a thumbs down. However, the session data it retrieved did not include the comment field, and it did not refer to the comment shown in the details panel.
When we asked for an analysis that combines 30 days of comments, it answered that this would be possible if the comment text could be retrieved, but not with the data it can currently access, and asked us to paste the comments.
Problem
The AgentOps agent can estimate causes and suggest fixes, but its inputs are incomplete.
Because end-user comments are not used in the estimation, it may identify a cause that does not match what the user described. Also, a person must find and specify the target conversations. In addition, it cannot group the causes shared by multiple pieces of feedback, or show how many conversations without feedback had the same problem.
Proposed Enhancement (in priority order)
Root solution: Enable the AgentOps agent to work end to end, from analyzing the cause to suggesting fixes, starting from end-user feedback. The goal is a semi-automated improvement loop, from feedback to identifying the cause, fixing, and verifying, while people keep the decision to adopt changes.
Retrieve feedback comments together with the rating and reason.
Collect conversations with feedback simply by specifying the agent and time range.
Compare the comments, conversations, traces (execution records shown in Trace Inspector), and the agent configuration at that time, and group the feedback by root cause, that is, the part of the configuration that caused the problem.
For each root cause, generate judge criteria and apply them with Custom LLM-as-a-Judge to conversations without feedback, to show the scale of impact. As with the prebuilt evaluators, the time range and sampling rate can be configured.
Suggest fixes. Use the existing AgentOps optimization for fixes to Instructions, and create test cases from the affected conversations for verification.
Configuration changes are not applied automatically. The people who operate the agent decide whether to adopt them.
Alternative: If the root solution is difficult, at minimum, enable the AgentOps agent to retrieve feedback comments (step 1). The AgentOps agent answered that if the comment text can be retrieved, it can produce a summary, frequent issues, the main causes of negative ratings, and suggested improvements.
Impact
This applies to production agents that collect end-user feedback in watsonx Orchestrate chat or the embedded web chat. In particular, where people other than the builder maintain the agent configuration, the effort to turn feedback into improvements is reduced. Other agent platforms, such as LangSmith Engine, have started to provide features that diagnose causes from production records and propose fixes. In watsonx Orchestrate, extending the existing AgentOps agent would provide a feedback-to-improvement loop within the product.
Related Ideas
LSABER-I-722 (Delivered): Collection and export of feedback. A prerequisite for this idea.
LSABER-I-1071 (Delivered): API for conversations and traces. The data source for this idea.
LSABER-I-924 (Planned for future release): Identifying the areas where the agent fails. This idea uses end-user feedback as the starting point and interprets it with the configuration and conversation records.
LSABER-I-1912 (Under review): Using feedback for reinforcement learning. This idea does not perform learning. It supports the human review that comes before such learning.
LSABER-I-2100 (Under review): Include the selected reasons and comments in the feedback export. It shows that there are currently limited ways to retrieve comments in bulk.