Sarthak Ahuja scite author profile

Skill routing is an important component in large-scale conversational systems. In contrast to traditional rule-based skill routing, state-ofthe-art systems use a model-based approach to enable natural conversations. To provide supervision signal required to train such models, ideas such as human annotation, replication of a rule-based system, relabeling based on user paraphrases, and bandit-based learning were suggested. However, these approaches: (a) do not scale in terms of the number of skills and skill on-boarding, (b) require a very costly expert annotation/rule-design, (c) introduce risks in the user experience with each model update.In this paper, we present a scalable self-learning approach to explore routing alternatives without causing abrupt policy changes that break the user experience, learn from the user interaction, and incrementally improve the routing via frequent model refreshes. To enable such robust frequent model updates, we suggest a simple and effective approach that ensures controlled policy updates for individual domains, followed by an off-policy evaluation for making deployment decisions without any need for lengthy A/B experimentation. We conduct various offline and online A/B experiments on a commercial large-scale conversational system to demonstrate the effectiveness of the proposed method in real-world production settings.

show abstract

Scalable and Robust Self-Learning for Skill Routing in Large-Scale Conversational AI Systems

Kachuee¹,

Nam²,

Ahuja³

et al. 2022

Preprint

View full text Add to dashboard Cite

Skill routing is an important component in large-scale conversational systems. In contrast to traditional rule-based skill routing, state-ofthe-art systems use a model-based approach to enable natural conversations. To provide supervision signal required to train such models, ideas such as human annotation, replication of a rule-based system, relabeling based on user paraphrases, and bandit-based learning were suggested. However, these approaches: (a) do not scale in terms of the number of skills and skill on-boarding, (b) require a very costly expert annotation/rule-design, (c) introduce risks in the user experience with each model update. In this paper, we present a scalable selflearning approach to explore routing alternatives without causing abrupt policy changes that break the user experience, learn from the user interaction, and incrementally improve the routing via frequent model refreshes. To enable such robust frequent model updates, we suggest a simple and effective approach that ensures controlled policy updates for individual domains, followed by an off-policy evaluation for making deployment decisions without any need for lengthy A/B experimentation. We conduct various offline and online A/B experiments on a commercial large-scale conversational system to demonstrate the effectiveness of the proposed method in real-world production settings.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Sarthak Ahuja

Examining the Effects of Anticipatory Robot Assistance on Human Decision Making

Modelling the sustainable supply chain management practices in Indian industries: A business model using Fuzzy TOPSIS approach

Competitive Coevolution for Color Image Steganography

Similarity Computation Exploiting the Semantic and Syntactic Inherent Structure Among Job Titles

Modelling the sustainable supply chain management practices in Indian industries: a business model using the fuzzy TOPSIS approach

Learning Vision-Based Physics Intuition Models for Non-Disruptive Object Extraction

Scalable and Robust Self-Learning for Skill Routing in Large-Scale Conversational AI Systems

Scalable and Robust Self-Learning for Skill Routing in Large-Scale Conversational AI Systems

Contact Info

Product

Resources

About