Bridging the Gap: Leveraging Crowdsourced Wisdom for Capstone Software Projects
Enriching Capstone Project-Based Learning Experiences Using a Crowdsourcing Recommender Engine
The paper introduces a Crowdsourcing Recommender Engine (CRE) integrated into an Educational Software Platform (ESP) to support Software Engineering (SE) capstone projects. By leveraging StackExchange data through BM25 and Latent Dirichlet Allocation (LDA), the system provides personalized resource recommendations based on real-time project activities.
Executive Summary
TL;DR: This research tackles the "information overload" problem in Software Engineering (SE) education by introducing a Crowdsourcing Recommender Engine (CRE). By analyzing student project data (tasks, commits) and matching it with StackExchange posts using LDA and BM25, the system provides high-precision, relevant technical resources at the moment students need them most.
Positioning: This work is a practical application of Information Retrieval (IR) techniques within the Computer Science Education (CSE) domain, evolving project management platforms into intelligent learning assistants.
The "Cold Start" of Project-Based Learning
Capstone projects are designed to mimic the software industry's pressure and complexity. However, students frequently encounter a wall: they must use languages, frameworks, or APIs they have never seen.
Traditional search engines often fail these students because:
- Context Misalignment: Search results aren't tailored to the specific stage or technical stack of a long-term project.
- Cognitive Load: Filtering through hundreds of StackOverflow threads is time-consuming and distracting.
The authors hypothesized that the information stored in project management tools (like Kanban boards and commit logs) contains enough "signal" to automate this discovery process.
Methodology: From Project Logs to Latent Topics
The researchers developed an architecture that transforms raw project metadata into actionable recommendations.
1. Data Aggregation and Cleaning
The engine extracts data from the Educational Software Platform (ESP), merging tasks, sub-tasks, and commit messages associated with a unique project ID. This text is then tokenized and cleaned.
2. Project Profiling with LDA
The core innovation lies in using Latent Dirichlet Allocation (LDA) to define the project's profile. Instead of just looking at word frequency (TF-IDF), LDA identifies the underlying topics.
- The engine identifies 10 topics per project.
- It selects the most representative words from these topics to form a "query vector."
3. Ranking with BM25
Unlike simple Cosine Similarity, which can be biased by document length, the engine uses Okapi BM25. This treats the project profile as a query and the StackExchange posts as the document collection, ranking them based on probabilistic relevance.
Figure 1: The architecture of the Crowdsourcing Recommender Engine (CRE).
Evidence of Success
The system was tested with students at the Pontificia Universidad Católica de Chile. The experiment compared four combinations of techniques.
Quantitative Performance
The results were clear: BM25 combined with LDA was the superior approach.
| Technique | Project Profile | Precision |
|---|---|---|
| TF-IDF + Cosine | TF-IDF | 0.229 |
| BM25 | LDA | 0.475 |
The precision nearly doubled compared to the baseline. This suggests that LDA's ability to "distill" the essence of a project into topics provides a much more effective query for the BM25 algorithm.
Qualitative Feedback
Students reported high levels of Novelty and Transparency. They understood why a post was recommended and often found resources they wouldn't have discovered through manual search.
Critical Insights & Future Outlook
The success of the CRE highlights a critical shift in educational tools: Context is King. By passively observing student behavior (via their project management tasks), we can provide active support.
Limitations:
- The current system relies on an MVP (Minimum Viable Product) and processes some data "on the fly," which may lead to latency issues as the dataset grows.
- The precision, while improved, still leaves room for growth (0.475 means roughly half are relevant).
Future Work: The authors aim to expand the knowledge base to include Reddit, StackOverflow, and even internal "crowdsourced" knowledge from past student groups. This "Institutional Memory" could prove invaluable for recurring challenges in capstone courses.
Conclusion
This paper demonstrates that the "saturation" of information on the web can be managed through intelligent recommendation. By marrying topic modeling with traditional information retrieval, we can transform a student's project board from a static list of tasks into a gateway to global expertise.
