From Search Logs to Silver Data: Automating Professional QA Systems
Exploiting Search Logs to Aid in Training and Automating Infrastructure for estion Answering in Professional Domains
This paper introduces a method to build Question Answering (QA) systems in professional legal and regulatory domains by leveraging "silver data"—Implicit Relevance Feedback (IRF) derived from search logs. It demonstrates that user interactions like saving or printing documents can substitute for expensive expert annotations, effectively addressing the cold start problem and boosting SOTA ranking performance.
TL;DR
In specialized domains like law, training a Question Answering (QA) system is notoriously expensive due to the need for expert-level "Gold Data." This paper from Thomson Reuters reveals a breakthrough: by mining "Silver Data" from user activity logs (actions like printing or saving), we can bypass the "Cold Start" problem. Their approach achieves performance comparable to having 25% of a full expert-labeled dataset, without the cost.
The Professional Data Gap: Why "Clicks" Aren't Enough
In general web search, a "click" is a noisy signal. In professional legal research, however, the stakes are higher. A lawyer doesn't just click—they save, print, or export a document when it solves a specific legal query.
The authors argue that these high-intent actions are a goldmine of implicit feedback. Before this study, the industry relied on Subject Matter Experts (SMEs) to manually grade thousands of document pairs—a process that is slow, unscalable, and creates a massive barrier for new entries (the "Cold Start" problem).
Methodology: The "Silver" and the "Imputed"
The research utilizes a two-stage retrieval architecture:
- Recall Stage: An internal legacy engine retrieves a broad set of candidates.
- Precision Stage (The Re-ranker): A learning-to-rank algorithm re-sorts these candidates based on features like n-gram overlap, embeddings, and parse tree correspondence.
Data Innovation
- Silver Data: QA pairs where users printed/saved/emailed the result. A validation study confirmed that 90% of these "Silver" documents were indeed relevant.
- Imputed Negatives: To teach the model what a "bad" answer looks like without asking an expert, the team sampled documents from deep within the search results (e.g., beyond rank 10), assuming they are highly unlikely to be relevant.
Table: SME validation confirming that Silver Data is 64% 'A' grade and 90% 'A' or 'C' grade.
Experimental Insights
The authors tested their hypothesis across two jurisdictions: Federal and State.
1. Solving the Cold Start
In scenarios where 0% gold data was available, the silver data provided an immediate, functional baseline. In the State jurisdiction, the performance leap was staggering—starting at nearly 50% accuracy at Rank 1, whereas sparse gold data only managed 16%.
2. Silver vs. Gold Efficiency
How much expert work does silver data replace?
- For Federal data, it took roughly 25% of the total gold data to match the performance of automated silver data.
- For State data, which is more structurally varied, silver data was even more dominant, outperforming small-to-medium gold datasets across the board.
Figure 1: Baseline performance using only expert-graded Gold Data. Note the low starting point and high variance.
Figure 3: Performance with Silver Data added. The "Cold Start" (0% Gold) begins at a much higher accuracy level (~0.47).
Critical Analysis & Conclusion
This work proves that professional AI doesn't always need a human-in-the-loop for every training sample. By treating user behavior as a proxy for expertise, companies can build "automated infrastructure"—systems that learn and improve every time a user saves a document.
Key Limitations:
- The "Imputed Negatives" strategy is simple; it doesn't specifically target "hard negatives" (documents that look relevant but aren't).
- Performance saturation occurs quickly as silver data grows, suggesting that while it's great for starting, high-end "Gold" data is still needed for peak optimization.
Future Outlook: This paves the way for Continuous Integration/Continuous Deployment (CI/CD) for AI models, where the system undergoes "Continuous Training" fed by real-time user logs, drastically reducing the time-to-market for legal and financial tech.
