Skip to Main Content
Shape the future of IBM watsonx Orchestrate

Start by searching and reviewing ideas others have posted, and add a comment (private if needed), vote, or subscribe to updates on them if they matter to you.

If you can't find what you are looking for, create a new idea:

  1. stick to one feature enhancement per idea

  2. add as much detail as possible, including use-case, examples & screenshots (put anything confidential in Hidden details field or a private comment)

  3. Explain business impact and timeline of project being affected

[For IBMers] Add customer/project name, details & timeline in Hidden details field or a private comment (only visible to you and the IBM product team).

This all helps to scope and prioritize your idea among many other good ones. Thank you for your feedback!

Specific links you will want to bookmark for future use
Learn more about IBM watsonx Orchestrate - Use this site to find out additional information and details about the product.
Welcome to the IBM Ideas Portal (https://www.ibm.com/ideas) - Use this site to find out additional information and details about the IBM Ideas process and statuses.
IBM Unified Ideas Portal (https://ideas.ibm.com) - Use this site to view all of your ideas, create new ideas for any IBM product, or search for ideas across all of IBM.
ideasibm@us.ibm.com - Use this email to suggest enhancements to the Ideas process or request help from IBM for submitting your Ideas.

Status Submitted
Created by Guest
Created on Oct 6, 2026

Enable the AgentOps agent to analyze causes starting from end-user feedback

Background / Current Behavior

In watsonx Orchestrate chat, end users can give a thumbs up or thumbs down to an agent response, select a reason, and optionally enter a comment. Currently, the details panel of the conversation view shows the number of positive and negative ratings and the comments for each conversation.

When we specified a conversation, the AgentOps agent estimated the cause of the negative rating from the conversation content and suggested fixes. It answered that it can suggest fixes for Instructions, descriptions of tools and collaborator agents, and knowledge, but can apply changes only to Instructions.

However, on October 6, 2026, we confirmed the following:

  • When we asked it to summarize thumbs-down comments for a specific agent and explain the causes, it could not identify the relevant conversations.

  • When we specified the conversation ID, it confirmed that the rating was a thumbs down. However, the session data it retrieved did not include the comment field, and it did not refer to the comment shown in the details panel.

  • When we asked for an analysis that combines 30 days of comments, it answered that this would be possible if the comment text could be retrieved, but not with the data it can currently access, and asked us to paste the comments.

Problem

The AgentOps agent can estimate causes and suggest fixes, but its inputs are incomplete.

Because end-user comments are not used in the estimation, it may identify a cause that does not match what the user described. Also, a person must find and specify the target conversations. In addition, it cannot group the causes shared by multiple pieces of feedback, or show how many conversations without feedback had the same problem.

Proposed Enhancement (in priority order)

Root solution: Enable the AgentOps agent to work end to end, from analyzing the cause to suggesting fixes, starting from end-user feedback. The goal is a semi-automated improvement loop, from feedback to identifying the cause, fixing, and verifying, while people keep the decision to adopt changes.

  1. Retrieve feedback comments together with the rating and reason.

  2. Collect conversations with feedback simply by specifying the agent and time range.

  3. Compare the comments, conversations, traces (execution records shown in Trace Inspector), and the agent configuration at that time, and group the feedback by root cause, that is, the part of the configuration that caused the problem.

  4. For each root cause, generate judge criteria and apply them with Custom LLM-as-a-Judge to conversations without feedback, to show the scale of impact. As with the prebuilt evaluators, the time range and sampling rate can be configured.

  5. Suggest fixes. Use the existing AgentOps optimization for fixes to Instructions, and create test cases from the affected conversations for verification.

Configuration changes are not applied automatically. The people who operate the agent decide whether to adopt them.

Alternative: If the root solution is difficult, at minimum, enable the AgentOps agent to retrieve feedback comments (step 1). The AgentOps agent answered that if the comment text can be retrieved, it can produce a summary, frequent issues, the main causes of negative ratings, and suggested improvements.

Impact

This applies to production agents that collect end-user feedback in watsonx Orchestrate chat or the embedded web chat. In particular, where people other than the builder maintain the agent configuration, the effort to turn feedback into improvements is reduced. Other agent platforms, such as LangSmith Engine, have started to provide features that diagnose causes from production records and propose fixes. In watsonx Orchestrate, extending the existing AgentOps agent would provide a feedback-to-improvement loop within the product.

Related Ideas

  • LSABER-I-722 (Delivered): Collection and export of feedback. A prerequisite for this idea.

  • LSABER-I-1071 (Delivered): API for conversations and traces. The data source for this idea.

  • LSABER-I-924 (Planned for future release): Identifying the areas where the agent fails. This idea uses end-user feedback as the starting point and interprets it with the configuration and conversation records.

  • LSABER-I-1912 (Under review): Using feedback for reinforcement learning. This idea does not perform learning. It supports the human review that comes before such learning.

  • LSABER-I-2100 (Under review): Include the selected reasons and comments in the feedback export. It shows that there are currently limited ways to retrieve comments in bulk.

Idea priority Medium