Cayley–Hamilton-Guided Krylov Regularization for Machine Learning: Theory, Computational Architecture, and Empirical Evidence
DOI:
https://doi.org/10.61467/2007.1558.2026.v17i5.1420Keywords:
Cayley–Hamilton theorem, matrix polynomials, Krylov subspaces, regularized least-squares classification, conjugate gradient, machine learning operators, Teorema de Cayley–Hamilton, subespacios de Krylov, gradiente conjugadoAbstract
This study develops and evaluates a Cayley–Hamilton-based framework for regularised machine learning. Its objective is to formulate a practical Krylov-subspace architecture for classification that replaces repeated dense inversion with stable operator iteration. The formulation draws on the theorem’s finite polynomial closure for square matrices. The analysis connects the characteristic-polynomial identity with regularised least-squares learning, derives full and truncated conjugate-gradient solvers, and evaluates the proposed method on four datasets. Full conjugate gradients reproduce the direct solution to machine precision, whereas truncated iteration preserves competitive predictive accuracy while reducing training cost and showing greater tolerance to perturbations in ill-conditioned data. These findings support the interpretation of the Cayley–Hamilton theorem not merely as an algebraic identity, but as a design principle for finite-dimensional learning operators, efficient computation and interpretable matrix-polynomial models.
Spanish-language metadata / Metadatos en español
Título en español:
Regularización de Krylov guiada por Cayley–Hamilton para aprendizaje automático:
Teoría, arquitectura computacional y evidencia empírica
Resumen:
Este estudio desarrolla y evalúa un marco basado en Cayley–Hamilton para el aprendizaje automático regularizado. Su objetivo es formular una arquitectura práctica de subespacios de Krylov para clasificación que sustituya la inversión densa repetida por una iteración estable de operadores. La formulación se fundamenta en el cierre polinómico finito del teorema para matrices cuadradas. El análisis vincula la identidad del polinomio característico con el aprendizaje regularizado mediante mínimos cuadrados, deriva solucionadores de gradiente conjugado completos y truncados, y evalúa el método propuesto en cuatro conjuntos de datos. Los gradientes conjugados completos reproducen la solución directa con precisión de máquina, mientras que la iteración truncada mantiene una precisión predictiva competitiva al tiempo que reduce el coste de entrenamiento y muestra una mayor tolerancia a perturbaciones en datos mal condicionados. Estos resultados respaldan la interpretación del teorema de Cayley–Hamilton no meramente como una identidad algebraica, sino como un principio de diseño para operadores de aprendizaje de dimensión finita, computación eficiente y modelos interpretables basados en polinomios matriciales.
Palabras Claves:
Teorema de Cayley–Hamilton, polinomios matriciales, subespacios de Krylov, clasificación regularizada por mínimos cuadrados, gradiente conjugado, operadores de aprendizaje automático.
Smart citations:
https://scite.ai/reports/10.61467/2007.1558.2026.v17i5.1420
Dimensions.
Open Alex.
References
Björck, Å. (1996). Numerical methods for least squares problems. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9781611971484
Engl, H. W., Hanke, M., & Neubauer, A. (1996). Regularization of inverse problems. Kluwer Academic Publishers. https://doi.org/10.1007/978-94-009-1740-8
Gazzola, S., & Sabaté Landman, M. (2020). Krylov methods for inverse problems: Surveying classical, and introducing new, algorithmic approaches. GAMM-Mitteilungen, 43(4), e202000017. https://doi.org/10.1002/gamm.202000017
Gazzola, S., Novati, P., & Russo, M. R. (2015). On Krylov projection methods and Tikhonov regularization. Electronic Transactions on Numerical Analysis, 44, 83–123. ETNA full text
Gohberg, I., Lancaster, P., & Rodman, L. (2009). Matrix polynomials. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9780898719024
Golub, G. H., Hansen, P. C., & O’Leary, D. P. (1999). Tikhonov regularization and total least squares. SIAM Journal on Matrix Analysis and Applications, 21(1), 185–194. https://doi.org/10.1137/S0895479897326432
Greenbaum, A. (1997). Iterative methods for solving linear systems. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9781611970937
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7
Hestenes, M. R., & Stiefel, E. (1952). Methods of conjugate gradients for solving linear systems. Journal of Research of the National Bureau of Standards, 49(6), 409–436. https://doi.org/10.6028/jres.049.044
Hoerl, A. E., & Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55–67. https://doi.org/10.1080/00401706.1970.10488634
Horn, R. A., & Johnson, C. R. (2012). Matrix analysis (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9781139020411
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. JMLR article
Qin, S. J., Liu, Y., & Tang, S. (2023). Partial least squares, steepest descent, and conjugate gradient for regularized predictive modeling. AIChE Journal, 69(4), e17992. https://doi.org/10.1002/aic.17992
Rifkin, R. M., & Klautau, A. (2004). In defense of one-vs-all classification. Journal of Machine Learning Research, 5, 101–141. JMLR article
Rifkin, R. M., Yeo, G. W., & Poggio, T. (2003). Regularized least-squares classification. In J. A. K. Suykens, G. Horvath, S. Basu, C. Micchelli, & J. Vandewalle (Eds.), Advances in learning theory: Methods, models and applications (pp. 131–154). IOS Press.
Saad, Y. (2003). Iterative methods for sparse linear systems (2nd ed.). Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9780898718003
Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427–437. https://doi.org/10.1016/j.ipm.2009.03.002
Suykens, J. A. K., & Vandewalle, J. (1999). Least squares support vector machine classifiers. Neural Processing Letters, 9(3), 293–300. https://doi.org/10.1023/A:1018628609742
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., . . . van Mulbregt, P. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Combinatorial Optimization Problems and Informatics

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.