Predictive Intelligence in Publishing: Solving the Print-Run Dilemma with Machine Learning
KNOWLEDGE‐BASED SYSTEMS
This paper presents a computational intelligence framework for predicting total sales of newly published books using real-world editorial data from Spain. It utilizes a three-stage approach—data visualization, feature selection (CFS, ReliefF, SOM), and machine learning regression (M5P trees, SVM, ELM, etc.)—to achieve high-accuracy pre-publication forecasting.
TL;DR
Publishing a new book is a high-stakes gamble. Print too many, and you lose money on pulped paper; print too few, and you lose money on missed sales. This research introduces a data-driven framework using real-world Spanish editorial data to predict book sales before they hit the shelves, achieving an impressive 0.89 correlation with actual outcomes using interpretable "Model Trees."
The Stakes: Inventory vs. Opportunity
In the world of physical books, the "Initial Print Run" is the most stressful decision a publisher makes. Existing literature often focuses on Time-Series Analysis, which is useless for a brand-new book with zero sales history.
Publishers have historically relied on "gut feeling" or subjective comparisons to "similar" titles. This paper argues that such an approach is unscalable and error-prone, proposing instead a Knowledge Discovery in Databases (KDD) approach to turn raw editorial logs into a structured forecasting engine.
Methodology: Beyond the Black Box
The authors didn't just dump data into a neural network. They followed a rigorous pipeline:
- Feature Engineering: They identified 11 key "pre-sales" variables, such as retail price, Dewey Decimal subject codes, and the number of intended points of sale.
- Dimensionality Reduction: Using Self-Organizing Maps (SOM) and ReliefF, they filtered out redundant data.
- Regression Comparison: They tested six algorithms, including Support Vector Machines (SVM), Multilayer Perceptrons (MLP), and the winner: M5P Model Trees.
Why M5P?
Unlike a standard Neural Network, the M5P algorithm produces a Decision Tree where the leaves are linear regression equations. This provides "Interpretability"—a publisher can look at the tree and see why the model suggests a 5,000-copy print run (e.g., "Because it’s a Fiction title priced under €20 with 500+ points of sale").

Key Insights from the Data
The study revealed striking patterns through SOM visualization. By looking at "Component Planes," the researchers found that 'Print Run' and 'Points of Sale' are almost perfectly correlated, suggesting that publishers' current strategies are heavily biased by their initial distribution capacity.

Performance Highlights:
- Accuracy: M5P and SVM led the pack, but M5P was significantly more efficient.
- Speed: While SVM took over 6,000 seconds to optimize, M5P finished in under 20 seconds.
- Interpretation: The model proved that Novelty Distribution (sending books to be featured on front tables) is a massive predictor of total sales, more so than the author's previous fame in some segments.

Case Study: The "Best Seller" Logic
The research applied the model to a "Top Seller" Spanish publisher. The resulting decision tree (pictured below) acts as a roadmap. For instance, if a book has over 1,006 points of sale and is distributed as a "Novelty," the baseline sales start at 24,000 units. If the retail price is too high, the model automatically penalizes the forecast.

Critical Analysis & Conclusion
Takeaway: This work proves that computational intelligence can outperform traditional editorial "expert knowledge." By focusing on variables under the publisher's control (price, distribution), the tool doesn't just predict the future—it helps shape it.
Limitations: The model is highly effective for "Mainstream" books but may struggle with "Black Swan" events (viral social media hits) since the dataset pre-dates the height of TikTok/Instagram influence (BookTok).
Future Work: The authors suggest integrating Fuzzy Logic to better handle the "vague" qualitative aspects of literature and moving toward real-time optimization where the print run is dynamically adjusted based on early "word-of-mouth" signals.
