Scorecard Receives Poor Grade

Smart Thinking: A Skill Versus Luck Essay Series | Issue 10

Written by Michael A. Ervolini

Scorecard Receives Poor Grade

INTRODUCTION

Ambivalence aptly describes how many investors feel about actively managed funds – especially equities. The fact that 50% of managed equities are now in passive funds highlights this point. Passive equity products are clearly the right solution for achieving a number of overall portfolio objectives (e.g., diversification, convenient exposures, parking cash short term, minimizing fees). Frequently, however, what motivates passive allocations is not a financial objective: Rather it is plain old fear.

The power of fear in shaping allocation decisions is enormous. It stems in good part from the weak information available about manager skill.[1] The angst generated by this informational void is then hypercharge by the gloomy reporting on the overall performance of active funds. The net effect is to position active equities so marginally as to render them less and less desirable as an asset class.

The idea of a “fear factor” effecting allocations to active funds is bolstered by newly published research. This work describes how the SPIVA Scorecard (Scorecard) persistently overstates underperformance among actively managed mutual funds. By doing so the Scorecard is authoritatively (I’ve certainly cited this source many times) misleading investors into believing that active equity management outcomes are far worse than reality. This overstatement of underperformance suggests that the Scorecard may be doing a disservice to investors and the market.

THE TITLE SAYS IT ALL

In their paper “How the SPIVA U.S. Scorecard Understates the Performance of Actively Managed Mutual Funds” authors K.J. Martijn Cremers, Jon Fulkerson, and Timothy Riley consider the impacts of several Scorecard assumptions on the report’s results. [2] The three Scorecard assumptions they investigate, and the alternatives assumptions used in their comparative analysis, are:

·  Methods to eliminate survivorship bias. The Scorecard assigns any fund that ceases to exist during an analysis period as underperforming in all years of the analysis. Cremers et al used the actual results from a discontinued fund during the years it was in operation.

·  Weighting results. The Scorecard weighted its results by number of funds, effectively giving an equal weight to all funds regardless of size. Cremers et al weighted their results by assets under management, essentially dollar-weighting their results.

·   Benchmark selection. The Scorecard uses a relatively small number of broad indices or what the authors refer to as hypothetical benchmarks (i.e., they cannot be invested in directly). Cremers et al compared each actively managed fund to the most closely matched passive product (a truly investable passive alternative).

After computing results based on their alternative assumptions the authors compared their levels of underperformance to those calculated by the Scorecard. They found that the results were consistently more negative for each individual Scorecard assumption versus their alternative assumption. The combined effect from all three assumptions they observed frequently produce completely opposite findings.

MANAGING ONE BIAS BY CREATING ANOTHER

When funds that are terminated during a multi-year analysis are excluded from a cohort investigation the potential exists for what is referred to as survivorship bias. The reasoning goes something like this: Funds that cease to operate at any time during the analysis period do so, it is assumed, because they are underperforming. It is further assumed that the funds that survive throughout the full period generated returns on average at least higher than those funds which were closed (if not actually outperforming their benchmarks). If the defunct funds are then excluded from the analysis the results are believed to be skewed more positively than had the excluded funds been included (i.e., results are biased toward the higher performing survivor funds). Although such performance tilt is not found among the survivors in every fund

cohort, safeguarding against survivorship bias is, nevertheless, a common analytic practice. So far so good.

The approach that the Scorecard takes to guard against survivorship bias is – well let’s just say ingenious. For any fund operating at the beginning of an analysis period and which then ceases to exist before the period ends, the Scorecard assumes: a) that from the day the fund closes until the end of the analysis period such fund remains in its cohort and is deemed to be underperforming, and b) during the period the fund did exist it is also designated as underperforming regardless of its actual performance. In other words, if a fund drops out at any time it is automatically assumed to be an underperforming fund from the first day of the analysis period on through to the final day. This method of correcting survivorship bias has the perverse effect of biasing the results toward greater underperformance. Moreover, the effect this method has on results increases monotonically over time. Because fewer funds survive year after year more and more funds are branded as underperforming as the analysis horizon lengthens.

Cremers et al recomputed the Scorecard results using what seems a more straightforward method for managing survivorship bias. Their alternative method simply includes all funds during the years in which they are in operation. Each fund’s performance is based on its actual returns and, importantly, none are transmuted into zombie nonperformers once they cease to exist. Unsurprisingly, greater levels of underperformance are reported by the Scorecard than those computed by Cremers et al, especially over longer periods. For example, between 2004 and 2024 (20 years) the percent of all equity funds reported as underperforming by the Scorecard and Cremers et al were approximately 94% and 78%, respectively (a 16% difference). Within the U.S. large cap funds cohort, the Scorecard reported the percent of underperforming funds by slightly more than 18% greater than what is computed by the researchers. Even greater differences in results are found among several fixed income and hedge funds categories. These differences reflect comparisons based on changing only a single assumption.

REPLACING ALL THREE ASSUMPTIONS

Cremers et al evaluated the entire mutual fund universe, comparing the results of the Scorecard to outcomes they generated using their three alternative assumptions. The comparisons produced impressive results even over relatively short time horizons, as the researchers explain: “Each change has a significant impact on the results. For example, the Scorecard method indicates that, over the 3-year period of 2022 through 2024, 81% of large-cap core funds underperform. Not automatically counting exiting funds as underperformers decreases that percentage to 77%. Likewise, weighting by assets decreases the underperformance to 71% and comparing against equivalent passive funds decreases it to 73%. Making all three changes simultaneously decreases the percentage to 46%. That is, the results invert, from a supermajority of active funds underperforming to a small majority of assets outperforming.”

Cremers et al found even greater differences over longer periods. As the researcher’s report: “Averaging across the cap- valuation U. S. equity categories, the Scorecard method indicates that 92% of funds have underperformed over the last 20 years. Among those categories, the lowest percentage [in the Scorecard] is only 86% (large- cap value). After our changes, we find an average of 55%, with half of the categories having an underperformance rate for their assets below 50%. Thus, while underperformance among U. S. equities remains the most common outcome after our changes, the likelihood approaches a coin flip.” While a 50/50 chance of the equity fund you select going on to outperform may not seem heartening it is far better than the paltry 8% that the Scorecard would have you believe.

MORE FINDINGS

Cremers et al found even greater differences across fixed income and hedge fund products. One example they cite regarding fixed income is: “Using the 10-year horizon (i.e., the longest horizon with full coverage), the Scorecard method indicates an across-category average underperformance rate of 71%. After our changes, that rate falls to just 37%, meaning that nearly two out of three dollars invested in active fixed income funds outperformed equivalent passive funds over the last 10 years.” They go on to observe: “Among high yield funds, one of the largest fixed income categories, the 15-year horizon underperformance rate decreases from 74% to 12% after our changes. Thus, within the fixed income class, our changes fully reverse the Scorecard’s conclusion: the data support the finding that active fixed income funds tend to outperform.”

BELEAGUERED PERHAPS, BUT NOT TO BE FORSAKEN

Active equity funds pose a challenging allocation dilemma. Keen assessment is required to identify funds more likely than not to outperform going forward. Such analysis is, however, severely hampered by the current lack of widely available rigorous manager skill metrics. Eschewing active equities altogether is an option that sidesteps the emotional challenges of selecting an actively managed fund. Unfortunately, it requires eliminating the diversification and excess return potential of an entire asset class. This is the uncomfortable choice facing investors today.

Actively managed funds actually perform significantly better than suggested by the Scorecard according to new research conducted by Cremers et al. The improved outcomes they identified are the result of substituting what are arguably three reasonable assumptions for those used in producing the Scorecard. It is unclear just what impact these findings will have on investor behavior. A hopeful speculation is that it may cause some investors to rethink their allocations to this asset class. While not the focus of their work, the insights Cremers et al provide underscore both the shortcomings of fund outcomes as proxies for skill and the urgent need to begin integrating decision-based skill metrics into active equity assessment.

ENDNOTES

  1. Michael A. Ervolini, “Skill Versus Luck – Taking The Guessing Out Of Equity Fund Selection,” MIT Press, 2026.
  2. K.J. Martijn Cremers, Jon Fulkerson, and Timothy Riley, “How the SPIVA U.S. Scorecard Understates the Performance of Actively Managed Mutual Funds”, May 4, 2026. Available at SSRN: https://ssrn.com/
Michael Ervolini headshot

MICHAEL A. ERVOLINI, AUTHOR

The ideas expressed on this website are developed and/or curated by Michael Ervolini. Mike has spent his entire 35 year + career leading efforts to improve and strengthen active management.

skill-versus-luck-logo-white

The multi-trillion dollar active management industry is predicated on the idea that managers have skill – yet little is known about it – Who has skill? How is it measured? This website is dedicated to finding answers to the questions surrounding skill.


SKILL VERSUS LUCK © COPYRIGHT 2026

Discover more from Skill Versus Luck

Subscribe now to keep reading and get access to the full archive.

Continue reading