AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Sk Repackages a Capacity Assumption

A ranking reversal driven by K, with a 2.4x dollar edge funded by tripled capital

2026-09-08 · 7 min read · US equities and ETFs

Reviewing: Knowledge-Optimising Investment Decisions with Informative Datasets · Sidharth Mallik and Waymond Rodgers · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Monthly Knowledge-Adjusted Multi-Configuration MPT Selector for US Equities

Backtest period 2015-01-01 to 2024-12-31 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

This is not a replication of the paper (2)

  • The paper does not define a universal observable knowledge-unit K; a backtest must choose an operational proxy such as number of datasets/models used, compute time, or assigned internal cost.
  • The empirical test would evaluate our chosen proxy and portfolio candidates, not a fully specified alpha signal from the paper.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (10)

  • deviation left undescribed by the audit (invalidates: The paper's scenario-only claims are now tested in a specific US equity/ETF sample and may not generalize outside this implemented universe.)
  • Universe (STOCK-only US large-cap top-250 vs paper's generic N investable securities): The paper is methodological with generic N securities; the spec trades a specific top-250 US-capitalization sample, now restricted to STOCK to eliminate the 9 near-duplicate ETF pairs (r=0.9919–0.9994) that made the eq(2) min-variance covariance rank-deficient. (invalidates: The paper's scenario-only claims are tested only in this specific US equity sample and do not generalize outside it; any ETF-inclusive interpretation of the universe is not supported.)
  • Covariance estimated from daily returns scaled to monthly (vs plain historical monthly moments): Covariance is estimated from ~756 daily close-to-close returns and scaled to a monthly base (×21) rather than from 36 monthly observations; this is required so the observation count exceeds the asset count (N up to 250) and the sample covariance is full rank. (invalidates: Any claim that the min-variance weights reflect the estimated monthly return moments as stated in the paper's plain-historical-moment MPT; the covariance now rests on a daily estimation base the paper does not specify.)
  • Ledoit-Wolf shrinkage + ridge diagonal fallback (ridge_diagonal_epsilon=1e-6): The spec regularizes the covariance with Ledoit-Wolf shrinkage and a ridge diagonal fallback; the paper's MPT rests on moments estimated from historical returns with no statistical regularization. (invalidates: The paper's Zero-Knowledge-Proof / ex-ante optimality argument that the post-optimization difference from equal weight is purely the tacit knowledge extracted from historical moments: regularized weights partly reflect the shrinkage prior, not the estimated moments.)

6 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2015Total 31.6%Dec 2024
Sharpe
0.11
Total Return
31.6%
Max Drawdown
-96.3%
CAGR
2.8%
Volatility
29.5%
Beta vs SPY
1.09
Trades
15,802

Sk leaves every portfolio choice unchanged when a desk's resource use is fixed. Rankings move only across configurations that consume different resources. The paper's sole reversal comes directly from K: P1 uses three knowledge units, while P2 uses one. Gross excess return then favors P2 by 2.4x in dollars, provided the cheaper portfolio can take substantially more capital. I would retain Sk as a reporting measure for research and compute expenditure. Table 1 gives little reason to use it as a portfolio-selection objective.

A disclosure comes first. The paper supplies no observable definition of the knowledge unit K. We therefore used counts of data and feature families for each configuration. Our ten-year run evaluates that proxy. It does not evaluate any signal specified by Mallik and Rodgers.

The proposal from Mallik and Rodgers

The paper begins with an Informative Dataset, defined as "Datasets that contain material information for investments while not representing the price of any instrument." Examples include climate science series, social media data and sentiment indicators. Under the normative route described by the authors, this dataset enters the return model. They illustrate the setup with a GARCH-X style formulation in which the dataset appears as an external variable, written schematically as ri ~ F(ri,t, mu, sigma, ID). Markowitz then supplies the weights, and the portfolio is scored against benchmark return rb through S = mu_d/sigma_d.

Their attribution objection is straightforward. S contains no term for either the data or the model. When a third-party dataset decays under this normative route, the damage appears as a return shortfall without an identified cause, "becoming a blind spot."

Knowledge Optimisation addresses that concern through a three-stage process. Stage one puts the dataset into Rodgers' Throughput Model as the information source. The dataset feeds the decision structure, instead of appearing solely in the pricing equation. Stage two gives Markowitz a new reading. Equal weights and purely codified inputs prevail before optimisation; afterward, the weights differ. The paper labels that difference tacit knowledge. Markowitz's ex-ante optimality argument attests to its existence and is presented as a Zero-Knowledge Proof.

Stage three combines the Sharpe ratio with Return on Knowledge, Rk = r/K. The resulting measure is Sk = mu_d/(sigma_d × K). Management chooses K: "a chosen measure of production, depending on how the management decides." Suggested candidates include desk size, money invested into the unit and compute cost.

The paper contains no dataset, sample period, universe, estimation or backtest. It advances no empirical claim. Its abstract describes the method plainly: utility is illustrated "Through scenario analysis," while the process "By design" increases the importance assigned to knowledge in investment decisions. Two stipulated scenarios carry the argument.

The first is persuasive. Investment I and compute cost C define a per-rebalancing-interval budget, Kmax = I/C. The investment is withheld if the search spends Kmax units without finding a parameter set expected to clear rtarget. The authors write: "This places a constraint on real-time performance, even when backtests would have shown that rebalancing was possible in the absence of resource constraints." That cleanly captures a genuine operational failure mode.

When does Sk change a choice?

For fixed K, Sk = S/K is monotone in S. A manager comparing target returns, or portfolios produced by one team with one data stack, receives exactly the ordering produced by S. I read Sk as a tool for capacity allocation and attribution.

The paper goes further, saying the measure "adds to performance analysis of investment portfolios and therefore improves portfolio selection." Yet its own example looks like a review of expenditure. Sk identifies an underperforming strategy when an unchanged risk-adjusted return requires a larger K. The paper says this "could be occurring due to a depreciation of 3rd-party sourced inputs such as data, model, or both." That belongs in a spend review.

Sk has no meaningful absolute level. Any rescaling of K also rescales the ratio, and management controls the definition of K. A unit conversion leaves rankings intact. Changing the measure may alter them. Headcount, dollars invested into the unit and GPU-hours would not necessarily produce the same 3-to-1 ratio. That 3-to-1 determines the paper's central example.

Table 1 is decided in advance

P1 has mu_d 10%, sigma_d 5% and K = 3. Its S = 2.0 and Sk = 2/3. P2 has mu_d 8%, sigma_d 5% and K = 1, producing S = 1.6 and Sk = 1.6.

The paper next interprets K as trading desks, each with maximum investment Ik. On that reading, P2 "can handle three times the maximum investment." Put a million dollars into P1 and it earns $100,000 over benchmark. Put $3M into P2 and it earns $240,000. The implied dollar advantage is 2.4x, from $240,000 over $100,000. The same value appears in the ratio between the two Sk figures, which the paper treats as confirmation. Its source is stated openly: "This scalability has implication for gross returns."

Reverse the identity and the outcome is predetermined. Both portfolios have sigma_d of 5%, so their Sk ratio equals (mu2/mu1) × (K1/K2). With capital tripled, the dollar ratio is 3 × mu2/mu1. The expressions are identical. The exercise therefore yields no finding specific to P2. Different values of sigma_d would break the equivalence.

The tripling enters as an assumption and returns as a dollar result.

Dollar risk also rises, without being priced by the paper. Using its own 5% sigma_d and the stated $1M and $3M allocations, my arithmetic gives P2 at $3M active risk of $150,000. P1 at $1M carries $50,000. On each unit of that risk, scaled-up P2 remains the 1.6 portfolio, while P1 remains the 2.0 portfolio.

Linear capacity without impact or crowding drives the result. Three replicas of a single configuration are precisely where assumptions of no impact and no crowding become hardest to defend. Table 1 therefore does not establish that the measure "improves portfolio selection." It supports a tighter proposition: the cheaper configuration wins when it can be tripled for free.

Neither scenario includes transaction costs, turnover or slippage. Compute cost C appears only as the search budget. The paper does cite research on complex transaction costs while arguing that the cost of knowledge units should enter selection.

The authors also acknowledge a literature that assigns MPT limited applicability. They set it aside because "many of the counterarguments" arise "from either ex-post results, or inaccurate moment estimates, both of which while holding significance could be avoided in an ex-ante theoretical setting." A theoretical paper can make that choice. Implementation makes it expensive, because mu_d and sigma_d still require estimation. The conclusion admits the deeper problem when it describes Knowledge Optimisation as optimal in the presence of codified and tacit knowledge, "somewhat counterintuitive when unknowns are involved."

What we built

Because the paper leaves K open, we created a proxy based on counts of data and feature families. Price-only received K=4, valuation-only K=4, quality-growth K=6, combined valuation-quality-momentum K=10 and text-enhanced K=12. We chose I=9 and C=1, giving Kmax=9. The two configurations with the heaviest data requirements were excluded immediately.

The paper reports no performance numbers. Our 31.56% and 0.11 Sharpe consequently have no figures from the authors beside them. These results describe our proxy and our construction, not their process.

We ran a monthly long-only selector on the top 250 US names by point-in-time capitalisation, excluding ADRs, from 2015-01-01 to 2024-12-31. Each candidate estimated expected one-month returns. Covariance used as many as 756 daily observations, with 504 minimum, and Ledoit-Wolf shrinkage. Candidates solved long-only portfolios across an active target grid from 0 to 1.00% monthly. Portfolios had a 10% weight cap and at least 30 names, after which each candidate retained its best feasible point.

Among the surviving candidates, we selected the one with the highest Sk. Cash was held when none had positive Sk. The benchmark was the equal-weight universe. We charged commission of four tenths of a cent per share and modelled no slippage.

The outcome was poor. From 2015-01-01 through 2024-12-31, the book returned 31.56% in total and suffered a 96.31% peak-to-trough drawdown. Sharpe was 0.11, Sortino 0.12, Calmar 0.03 and volatility 29.47%. A ten-year total return of 31.56% paired with a 96% drawdown is a bad result from our construction. We assembled it from the authors' idea. It is one implementation of our choices and does not test a measured claim by the paper, which makes no measured claim.

Two features of our construction deserve direct treatment. The two K=4 candidates have identical rankings under S and Sk. The penalty mattered only by requiring quality-growth to exceed them by more than 50% on active Sharpe before selection. For our implementation, Sk became a 1.5x hurdle combined with hard exclusion.

The inputs matter as well. Values for mu_d and sigma_d came from a 36-month rolling return model and shrunk daily covariance. The paper sets aside exactly this estimation noise. Our 0.11 Sharpe and 96.31% drawdown belong to those choices, including zero modelled slippage for a monthly-rebalanced 250-name portfolio.

We have previously written about an index that ultimately tied the simpler measure it was designed to beat (/articles/the-triadic-stress-index-ties-the-measure-it-set-out-to-beat). Sk differs because it genuinely changes rankings. A denominator selected by the firm controls those changes.

One result would alter my view: an Sk based on a stated, auditable K that predicts which configuration preserves its active return at three times the capital. Until such evidence appears, the strongest part of the paper remains Kmax = I/C and the decision to withhold the investment.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each monthly rebalance date at the close:
  1. Form the point-in-time universe:
     - US STOCK instruments only
     - non-ADR filter
     - top 250 by PIT capitalization from aiquant_screening_table for the rebalance year

  2. Compute PIT features for eligible names:
     - 12-1 momentum: close[t-21] / close[t-252] - 1
     - 1-month reversal: close[t] / close[t-21] - 1, used with negative direction
     - 63-day dollar-volume liquidity and realized volatility
     - daily valuation metrics with date <= signal date
     - financial-statement quality, growth, leverage, cash-flow, and liquidity metrics with datekey <= signal date
     - winsorize cross-section at 1%/99%, z-score by month, median-impute analysis features only

  3. For each candidate portfolio:
     - skip if K > Kmax or required data/optimization inputs are insufficient
     - build expected one-month security returns from the candidate feature model
     - estimate covariance from up to ~756 daily returns, require at least 504 observations, scale by 21 to monthly, apply Ledoit-Wolf shrinkage plus ridge fallback
     - solve long-only MPT portfolios over the active target-return grid [0, 0.25%, 0.50%, 0.75%, 1.00%]
     - constraints: sum weights = 1, 0 <= weight <= 10%, minimum 30 valid assets
     - retain the feasible target-grid solution with the highest estimated standard active Sharpe S = μd/σd
     - compute Sk = μd / (σd × K) and require Sk > 0

  4. Select the candidate with the highest Sk among feasible K <= Kmax candidates.
     - comparison selector logs the highest ordinary active Sharpe S without K penalty
     - if no candidate passes, withhold allocation and hold cash

  5. Rebalance at the close using real close execution prices only.
     - skip trades with missing execution prices
     - apply commissions; no slippage is modeled