Every morning at 06:30 the mill ingests new finance papers from arXiv, SSRN, journals and research blogs, screens them for relevance, and scores the survivors against a robustness rubric: edge size, buildability with data we can actually get, methodology hygiene, breadth, novelty against everything we have already killed. Papers that pass get a house test with our own data, execution assumptions and costs. Papers that fail get a reason.
One change from issue 1. Every strategy paper below leads with the gross number the authors report, followed by the cost per side that would take it to zero, which is our arithmetic from the paper's holding period and turnover. A retail account paying 20 basis points and a desk paying 2 read the same paper differently, and a thin net edge at our costs can be a fat one at yours. Verdicts on cost are yours to make; verdicts on the evidence are ours.
This issue covers two weeks, because last Thursday's did not go out; the scoring pipeline stalled on an exhausted API balance, which is now fixed. Numbers below are the authors' unless marked as ours.
Snapshot covering records ingested 2026-09-13 to 2026-09-25: 2,607 papers ingested (most from a bulk topic sweep, and most rejected at the abstract stage or waiting for full text), 230 with recorded scores and 1 marked promoted to the bench. Separately, 8 briefs were published 2026-09-13 to 2026-09-25. The seven below are the ones I would spend twenty minutes on.
1. Employees see it first: Glassdoor rating changes and the next quarter
T. Clifton Green, Ruoyan Huang, Quan Wen and Dexin Zhou, "Crowdsourced Employer Reviews and Stock Returns", Journal of Financial Economics 2019 (SSRN 3002707).
Gross: 0.74% a month, value-weighted, four-factor alpha 0.77% a month. Sort firms with at least 15 Glassdoor reviews a quarter on the change in their average rating, June 2008 to June 2016, 1,238 firms. Long the top quintile of improvers, short the bottom quintile, hold three months, rebalance quarterly. Equal-weighted, the long-short alpha is 0.88% a month (t = 2.70); the long-short Sharpe ratio is 0.98. The signal is in current employees' reviews (0.88%), not former employees' (0.28%, insignificant), and in the career-opportunities and senior-management ratings, not work-life balance. Rating changes predict sales growth, profitability and the next quarter's earnings surprise, which is the mechanism: employees see the business turn before the market does.
Break-even cost (ours): about 55 basis points a side. Quarterly rebalancing with full turnover of both legs is four sides a quarter against roughly 2.2% of gross spread, so the edge survives any cost a liquid-name investor actually pays. The paper reports no costs, and it does not need to.
Mill priority: Test. The questions are not about cost. The sample ends in 2016, and Glassdoor data has been sold to systematic investors since, so the first thing to check is whether the effect survived its own publication. The second is point-in-time discipline: reviews are stamped on submission, and the gap to public visibility is not documented.
2. Skewness-managed anomalies
Rui Gong, John Lynch and Richard Ogden, "Skewness Managed Portfolios" (SSRN 6913978).
Gross: +5.45 percentage points a year, on average, across 18 anomalies; Sharpe ratios up 0.12. Many anomaly returns come from a handful of extreme, positively skewed stock returns. The authors predict each stock's skewness from characteristics, then tilt each anomaly toward high-skewness names in the long leg and low-skewness names in the short leg. CRSP common stocks, July 1963 to December 2024. The largest gains are in Value and Investment, the gains concentrate in periods of stress, and the managed portfolios keep significant alpha even against factor models built from the same characteristics.
Break-even cost (ours): about 11 basis points a side on the increment alone, if the tilt turned over both legs in full every month. That is the worst case; a within-leg tercile sort on a monthly rebalance turns over a fraction of the book, so the true break-even is higher. The paper says the gains survive the Novy-Marx and Velikov cost model but prints no turnover, which is the number a reader needs.
Mill priority: Watch. The sign runs the other way from the older idiosyncratic-skewness literature, where lottery-like stocks earn less, and the paper does not explain why total skewness should pay when idiosyncratic skewness is penalised. The 45 basis points a month is real money if the turnover is modest. Show the turnover and this moves to Test.
3. A Kalman filter and a controller, judged on drawdown
Yu Peng, Matloob Khushi and Josiah Poon, "CAST: A Cross-Asset State-Space Trading System for Drawdown Control in Stock Markets" (arXiv 2609.14205, accepted at ICDM 2026).
Gross: $1,000 to $1,633 over 2010 to 2025 on a 30-stock NASDAQ basket, maximum drawdown 11.4%, annualized Sharpe 0.52. A cross-asset Kalman filter estimates each stock's latent state online; a model-predictive controller turns the forecast into trades with forecast uncertainty as an explicit risk penalty. Calibrated on 2005 to 2010 and left untouched for fifteen years. On a five-currency global basket, $1,133 with a 3.1% drawdown. In the 2020 crash the NASDAQ basket drew down 7.1% against 34.0% for the market. On the authors' table CAST holds the return-drawdown frontier on every market.
Break-even cost (ours): cannot be computed, because turnover is not reported and the backtest is explicitly frictionless. The controller re-solves at every close; at daily re-optimisation a single basis point a side could matter against 3.3% a year.
Mill priority: Pass as a strategy, watch as a control framework. The drawdown numbers are real engineering, and the modular design means the controller can sit on top of a forecast you already trust. What it will not do is generate the forecast.
4. Diffusion models for the volatility surface, judged by the hedge
Yinbin Han, Jack Yuxiang Zhang, Manuel Torres, Fernando Acero and Renyuan Xu, "Diffusion models for dynamic volatility surface generation and data-driven hedging" (arXiv 2609.13402).
Gross: no P&L is claimed; the result is variance reduction. A diffusion model learns the joint next-day evolution of the S&P 500 return and its implied-volatility surface from daily SPX option data, 2000 to 2023, and a fine-tuned version pushes static-arbitrage violations to almost zero. Generated scenarios feed a hedge optimiser. Pooled across strikes, the unhedged position has a tracking-error standard deviation of $118.95, delta hedging $41.09, delta-vega $15.53, and the diffusion hedge lower still. Against the GAN benchmark at 0.9 moneyness, standard deviation falls from $68.66 to $10.49 and the 1% tail loss from $80.78 to $29.36, with the 2020 disruption included. Code is public.
Break-even cost: not applicable. Rebalancing cost is a penalty inside the optimiser, not a hurdle outside it.
Mill priority: Pass for trading, watch for risk. If you hedge options, the question it answers is whether a learned conditional distribution beats Greeks. On this data it does.
5. Commodity seasonality, tested until it broke
Ralph Kosch and Robin Forsberg, "Seasonal Trading in Commodity Futures: Evidence from Regression and Singular Spectrum Signals" (arXiv 2609.12227).
Net after costs: across 500 bootstrap paths for 2016 to 2024, classical SSA has the strongest average seasonal-model results: a 12.96% cumulative return and a Sharpe ratio of 0.131. Its median path loses 11.30%. Three seasonal models on 15 liquid commodity futures, estimated on rolling ten-year windows, with 8.6 basis points one-way costs already included and uncertainty from 500 maximum-entropy bootstrap paths. The comparison, an equal-weight long basket, does better on average: 16.81% cumulative, Sharpe 0.191, maximum drawdown 41.4%. None of the 18 paired Sharpe tests rejects after a within-family Holm adjustment.
Break-even cost: not the issue. The authors already charge 8.6 basis points a side. The benchmark pays the same monthly round-trip cost convention and has the stronger average bootstrap result. A zero-cost comparison is not reported here.
Mill priority: Reject, and read it anyway. This is what a null result should look like: a cost line, a do-nothing benchmark, a multiple-comparison correction, and a mean-versus-median check that catches the lottery payoff. The authors also say their bootstrap keeps the calendar fixed, so it is not a test against a timing-randomised null. Seasonal commodity signals are not dead in this paper. Across these bootstrap averages, they do not beat holding the basket; the historical Sharpe comparisons are statistically inconclusive.
6. When liquidity evaporates, it evaporates together
Demetrio Lacava and Paolo Santucci de Magistris, "Illiquidity at Risk" (arXiv 2609.00943).
Gross: no trade; a risk measure. Illiquidity-at-Risk is a quantile forecast of tomorrow's realized Amihud measure (realized volatility over volume), fitted with HAR and multiplicative-error models with and without a jump component. Data: the S&P 500 from January 2005 to October 2021 and 25 large US stocks from January 2012 to January 2024, with out-of-sample density tests. Models without jumps systematically under-cover the extreme events; adding the jump component corrects the coverage in stress periods. Single-stock violations cluster on days when the index's own liquidity is stressed, so the index leads the names, and the response is asymmetric in the same way volatility is: liquidity dries up faster after losses.
Break-even cost: not applicable.
Mill priority: Pass for trading, adopt as a monitor. The useful claim is that a cheap daily series (5-minute realized volatility over volume) forecasts the days on which your cost model is wrong. If your execution assumptions are constant, this is the paper that says when they are not.
7. A robustness score that admits it does not predict
Maria Laura Santoni, Vincent Jouanne and Matthew L. Scullin, "Equity Strategy Backtesting: Luck or Edge? The MinervaScore as a Statistical Robustness Grade" (arXiv 2608.23808).
Gross: none claimed; the paper's own pre-registered real-market test of 352 strategies found no forward relationship between score and out-of-sample Sharpe (Spearman 0.013, p = 0.40). The score combines the Deflated Sharpe Ratio, Probability of Backtest Overfitting, Superior Predictive Ability, Minimum Track Record Length and a regime-stability check into a 0 to 100 grade, gated so that a score of 80 or more appears only when all five pass. Calibrated on 359,062 production backtests, of which 3,848 (1.07%) pass every gate. In synthetic markets with known ground truth it separates real edge from lucky backtests with an AUROC of 0.989, almost identical to the Deflated Sharpe Ratio alone at 0.988. In the calibration population, 60% to 95% of Deflated Sharpe values are reported as exactly zero, depending on strategy family.
Break-even cost: not applicable.
Mill priority: Pass as a tool, keep the two numbers. The composite adds almost nothing over the Deflated Sharpe Ratio, which the authors say plainly. Two figures are worth remembering. In production, 1 in 93 backtests survives the five-gate battery. In the forward test, none of the 352 enrolled strategies passed all five gates, the median out-of-sample Sharpe ratio was minus 0.15 and only 45% were positive, so the score had almost no surviving edge to find, and the authors say so rather than claiming otherwise. Whether a strategy that does pass all five gates goes on to earn its backtest is untested here. Reported in full, with a null, by people selling the score. That is rarer than the score.
Briefs published in these two weeks
Eight pieces went out. Three research reviews on September 13: whether VIX improves a five-ETF monthly rotation, whether slowing down rescues futures trend following after modelled costs, and when trend following should yield to buy-and-hold. Then four house tests: Bitcoin's rising equity beta on daily data (September 15), the 162% intraday reversal paper against a 15 basis point spread hurdle (September 19), a machine-learning regime indicator on eight ETFs (September 20), and a free replication of the overnight-versus-intraday return split, which reproduces the famous chart and shows the US intraday loss ended in 2010 (September 21). On September 22, the constrained volatility-managed portfolio test: a sixth less volatility, and the risk-adjusted return did not improve. Links at financepapermill.com/research.
On the bench
One newly promoted paper is in protocol; we will report it when the test is complete, not before. Three of the window's higher-scoring papers overlap work in progress on our own desk and are held back from this issue; we publish tests, not previews. Last issue's leader, the end-to-end neural shrinkage of correlation matrices for small-cap portfolios (mill score 76), is still not scheduled.
Graveyard
Brogaard, Han and Kim's intraday residual-reversal paper reports 162% annualized before costs at midpoint prices. Our four-characteristic, high-VIX, two-hour adaptation averaged 4.38 basis points gross per long-short formation across the eligible portion of a 2017 to May 2026 sample; random selection averaged minus 0.32. Break-even, in this issue's format, is about 1.1 basis points a side; a fixed 3.86 basis points on each of four legs takes the mean to minus 11.06. The cost line was calibrated from a broader sample of entry quotes rather than matched executable fills, so the next step would need historically correct index membership and an actual bid-ask fill study. Mill verdict: failed the modelled cost gate at our costs; at 1 basis point a side it is a different conversation. Full house analysis: https://financepapermill.substack.com/p/the-162-intraday-reversal-paper-44
Forward this to one person who reads papers. The Mill Run is free. Paid subscribers get the house tests on Tuesdays.
Research findings are not a recommendation to trade. AI assists curation and drafting; human judgment selects the papers and signs every verdict.
Read the complete free Mill Run on Substack →
For measured house results, explore our house tests. Read how we evaluate the evidence.
Curated by a 30-year hedge fund veteran. AI assists curation and research preparation; human judgment selects the papers.