Watson Speech to Text should return timestamps accurate to milliseconds for transcription

Real-life scenario:

Researchers from Brandeis University, Boston University, Harvard, Boston College, and Northeastern University are investigating cognitive aging and biomarkers of dementia. Currently, they have been hand-scoring cognitive interviews which is arduous. To automate some of the manual work, Watson Text to Speech service is being used to run audio transcriptions through the service while getting back transcriptions with timestamps of when a word starts to be spoken and when it is ended in speech. The timestamps returned from the service return with 2 decimal places which is not enough precision that is required to compare to prior work (which uses 3 decimal places -- precision is up to milliseconds).

Problem statement:

The issue that researchers mentioned above are running into is comparing transcription timestamps generated by Speech to Text service to hand-scoring done for interviews prior to using Watson. The precision mismatch does not allow the researchers to use Watson effectively.

Current workaround:

No workaround is possible since there is a mismatch in precision level returned by Watson Speech-to-text service.

Proposed solution:

Having timestamps returned with 3 decimal places (with millisecond precision) would enable the researchers to automate the longitudinal research with Watson Speech-to-Text service.

Benefits/Value:

Watson Speech-to-Text service returning timestamps with milliseconds precision will enable customers around the world to get more accurate transcriptions and use the service more confidently in research globally. With precision, researchers will be more likely to use the service and cite it in publications.

Users impacted:

Every user using Watson Speech-to-Text around the world will get more precise data points for transcription without their current downstream applications breaking due to this change.

Idea priority

Urgent

Post comment

Admin

Marco Noel

May 9, 2022

I'm not sure I understand the business value of adding an extra decimal to a timestamp. Please expand.
Once we have more details, we will review but at this time, we are dealing with higher priorities.

Reply
Hide replies

Guest

Mar 3, 2022

Could I get an update on this please, this is a high priority request which will determine if the research will keep using IBM Watson or switch to a different Speech to text service.

Reply
Hide replies

By clicking the "Post Comment" or "Add Idea" button, you are agreeing to the IBM Ideas Portal Terms of Use.
Do not include IBM confidential, company confidential, or personal information in any field.
Having problems accessing this portal? Describe the problems in an email to ideasibm@us.ibm.com.

Please enter your email address

RELATED IDEAS