Research update · CEC-2013 LSGO, D=1000

Subspace optimization is a specialist, not a general improvement

Measured against plain Differential Evolution, the dual-EA subspace method loses on most problem classes and wins by roughly 50× on overlapping structure. That pattern is the finding. It tells us which problems the method is actually for.

Where it helps

Every run below is 3 million evaluations at D=1000. The parent is plain DE with no projection and static parameters (F=0.5, CR=0.9); the subspace arm is the dual-EA with LoRA rank 1, same optimizer, same budget, same seeds.

gain over plain DE by problem structure
Each function, ordered by structural class. Right of the dashed line the subspace method wins.
FunctionStructureplain DEsubspacegainseeds
F1fully separable (alignment artifact)3.439e+035.309e-066.5e+08x5
F13overlapping5.027e+091.045e+0848.1x5
F12single group8.324e+032.163e+033.85x1
F4partially separable1.910e+106.848e+092.79x1
F3fully separable3.936e+001.565e+002.51x1
F2fully separable4.794e+032.682e+031.79x1
F5partially separable1.011e+062.640e+060.383x5
F9non-separable (20 grp)7.393e+072.931e+080.252x5
F15single group3.718e+075.552e+080.067x5
F6partially separable4.030e+015.610e+030.00718x1

Aggregated by structural class

Classgeo-mean gainfunctions
overlapping48.12x1
fully separable2.12x2
single group0.51x2
non-separable (20 grp)0.25x1
partially separable0.20x3

The overlapping functions are the only class where the method wins. Overlapping problems share variables between subcomponents, so improving one component requires moving several together. That is the kind of coordinated move a low-rank step can make and a coordinate-wise local search cannot.

Is it this map, or any map of the same size?

All three projections searched at the same 64 numbers, same dual-EA scheme, same budget and seeds. If any reduction of that size performed alike, the structure of the map would be irrelevant and there would be nothing to study.

FunctionLoRA r=1random projectionrandom blockingseeds
F14.139e-068.748e+068.383e+062

The map matters: at identical search dimension the three are not interchangeable.

Convergence

convergence plot
DE, self-adaptive parameters. Fixed projections (solid orange, solid green) flatten early and stop improving.
convergence plot
DE with static parameters. This is the parent baseline we compare EvoSubspaceOptimization against.

Does self-adapting the optimizer's parameters help?

PyMOO can evolve DE's F and CR per individual. Matched pairs, same seeds and budget, static against self-adaptive:

ArmOptimizerpairsevolving winsgeo-meanverdict
no projectionDE269/260.45xstatic better
no projectionPSO1514/152.81xevolving better
EvoSubspaceOptimizationDE2617/2625.42xevolving better
EvoSubspaceOptimizationPSO2525/255.48xevolving better

It depends entirely on what you are adapting. Self-adaptation hurts plain DE, whose F=0.5, CR=0.9 defaults are already well tuned for this benchmark. It helps the dual-EA substantially. That arm runs two populations in very different regimes, and no single setting suits both.

Does learning a projection parameter help?

We made the projection's output scale a searched decision variable. That costs one extra dimension, and the neutral value reproduces the unscaled map exactly. It addresses a defect we measured: LoRA's factors overshoot the box by about 100×, which clips roughly 95% of coordinates.

ProjectionOptimizerpairsscaled winsgeo-mean
random projectionDE2525/252.04x
random projectionPSO2525/251.89x
LoRA r=1DE2520/252.31x
LoRA r=1PSO2518/251.64x
EvoSubspaceOptimizationDE2515/251.29x
EvoSubspaceOptimizationPSO2513/252.70x

The gain tracks how much work the projection is doing. Random projection, which must produce the whole solution, improves in every pair. The dual-EA, which is anchored on the full-space best and so already has its reach handled, barely moves.

Is the headline result real? A permutation control

Shuffling the variable order gives back the same function, with the same weights and the same difficulty. All it breaks is any accidental alignment between the benchmark's coordinate ordering and the method's internal reshape. We ran it over all 15 functions and 3 seeds:

Functioncost of shufflingeffect
F11.8e+12xcollapses; all of it was ordering
F75.50xordering irrelevant
F21.29xordering irrelevant
F81.23xordering irrelevant
F41.21xordering irrelevant
F131.04xordering irrelevant
F101.01xordering irrelevant
F111.00xordering irrelevant
F61.00xordering irrelevant
F50.93xordering irrelevant
F30.92xordering irrelevant
F120.89xordering irrelevant
F140.85xordering irrelevant
F90.71xordering irrelevant
F150.06xordering irrelevant

F1 collapses by about 1012. Nothing else moves. The median across the other fourteen is 0.998, so the typical function's result moves by a fraction of a percent. One function in fifteen was affected, and it happened to be the headline one.

Optimising the optimizer

A natural extension of the two-population idea is to let one population optimise the optimizer's own parameters rather than more solutions. We implemented that: a second population whose genome is (F, CR), scored by the progress the solution population actually makes under those values.

Variantrunsgeo-mean best
over fullspace202.700e+06
over dual-EA93.328e+05

It does not pay at this budget. Racing λ candidate parameter sets to advance one costs a λ-fold tax on progress, and the better parameters it finds do not repay that. Over plain DE it landed behind static parameters. Over the dual-EA it beat the evolving baseline in only 3 of 9 runs. We are reporting it as a negative result.

Does it hold at ten times the dimension?

One seed at D=10000, DE. Note this is a different problem family rather than a harder version of the same one: at D≠1000 the benchmark generates its structure from a seed instead of loading the official data files, so these are comparable to each other but not to the D=1000 numbers.

Functionplain DELoRArandom proj.EvoSubspaceOptEvoSubspaceOpt + scale
F11.674e+092.051e+122.140e+123.886e+081.433e+08
F56.629e+069.001e+069.848e+065.010e+066.253e+06
F136.571e+093.209e+201.337e+17not runnot run
F154.841e+092.338e+121.880e+184.367e+112.527e+10

The ordering is unchanged: the dual-EA variants lead, fixed projections trail badly.

What the controls rule out

F1 is an artifact. Its weights factor as 10^(c(32i+j)) = 10^(32ci) × 10^(cj), which is exactly the multiplication table the LoRA map produces. It is rank-1 to machine precision. A permutation control over all 15 functions and 3 seeds showed F1 collapsing by ~1012 when variables are shuffled, while nothing else moved. F1 is therefore best treated as a special case rather than evidence of a general effect.

Fixed projections are worse than no projection. Searching a fixed LoRA or random projection in absolute mode lands five to six orders behind plain DE. Clipping to the box turns a rank-4 solution into a full-rank saturated point, destroying the structure the method depends on.

Optimising the optimizer did not pay at this budget. Racing candidate parameter sets costs more evaluations than the better parameters return. Worth revisiting at larger budgets or with a cheaper credit-assignment scheme.