<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "https://jats.nlm.nih.gov/publishing/1.3/JATS-journalpublishing1-3.dtd"><article xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" dtd-version="1.3" article-type="research-article"><front><journal-meta><journal-id journal-id-type="issn">2460-0245</journal-id><journal-title-group><journal-title>Journal of the Indonesian Mathematical Society</journal-title><abbrev-journal-title>JIMS</abbrev-journal-title></journal-title-group><issn pub-type="epub">2460-0245</issn><issn pub-type="ppub">2086-8952</issn><publisher><publisher-name>IndoMS</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.22342/jims.v32i2.1960</article-id><article-categories></article-categories><title-group><article-title>Stochastic Influence on Levenberg-Marquardt Method for Nonlinear Least Squares Problems</article-title></title-group><contrib-group><contrib contrib-type="author"><name><surname>Christopher</surname><given-names>Gerend</given-names></name><address><country country="ID">Indonesia</country><email>gerendc@gmail.com</email></address><xref ref-type="aff" rid="AFF-1"></xref><xref ref-type="corresp" rid="cor-0"></xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-1789-7187</contrib-id><name><surname>Naiborhu</surname><given-names>Janson</given-names></name><address><country country="ID">Indonesia</country><email>janson@itb.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref></contrib></contrib-group><contrib-group><contrib contrib-type="editor"><name><surname>Rizal</surname><given-names>Jose</given-names></name><address><email>jrizal04@unib.ac.id</email></address></contrib></contrib-group><aff id="AFF-1"><institution content-type="dept">Department of Mathematics</institution><institution-wrap><institution>Bandung Institute of Technology</institution><institution-id institution-id-type="ror">https://ror.org/00apj8t60</institution-id></institution-wrap><country country="ID">Indonesia</country></aff><author-notes><corresp id="cor-0">Corresponding author: Gerend Christopher. Email: <email>gerendc@gmail.com</email></corresp></author-notes><pub-date date-type="pub" iso-8601-date="2026-04-23" publication-format="electronic"><day>23</day><month>04</month><year>2026</year></pub-date><pub-date date-type="collection" iso-8601-date="2026-04-23" publication-format="electronic"><day>23</day><month>04</month><year>2026</year></pub-date><volume>32</volume><issue>2</issue><issue-title>JUNE</issue-title><fpage>1</fpage><lpage>13</lpage><history><date date-type="received" iso-8601-date="2025-03-03"><day>03</day><month>03</month><year>2025</year></date><date date-type="accepted" iso-8601-date="2025-11-29"><day>29</day><month>11</month><year>2025</year></date></history><permissions><copyright-statement>Copyright (c) 2026 Journal of the Indonesian Mathematical Society</copyright-statement><copyright-year>2026</copyright-year><copyright-holder>Journal of the Indonesian Mathematical Society</copyright-holder><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/"><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-nc-nd/4.0/</ali:license_ref><license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p></license></permissions><self-uri xlink:href="https://jims-a.org/index.php/jimsa/article/view/1960" xlink:title="1960"></self-uri><abstract><p>In this digital era, a lot of data can be collected to be extracted and used for decision making through mathematical models in solving cases in certain domains. This paper will focus on solving nonlinear least squares problems in building an optimal model, namely one that has good accuracy and is efficient. In practice, the model is built using numerical methods. The main objective of this study is to investigate the effect of stochasticity in numerical methods that utilize gradients, namely Levenberg-Marquardt, on accuracy and computational efficiency. In addition, several numerical results from several variants of Stochastic Levenberg-Marquardt sampling and data taken in clusters with K-Means will be compared. The results of this paper are Stochastic Levenberg-Marquardt with 100, 200, 300, 400, and 500 data sample sizes outperformed classical Levenberg-Marquardt since they have very low relative errors to Levenberg-Marquardt and are faster in computational time than Levenberg-Marquardt. Therefore, when solving nonlinear least squares problems with large data sets, the Levenberg-Marquardt method requires only a small number of samples, which suggests that using the Stochastic Levenberg-Marquardt approach is advantageous.</p></abstract><kwd-group><kwd>nonlinear least squares</kwd><kwd>stochastic Levenberg-Marquardt</kwd><kwd>cluster</kwd><kwd>accuracy</kwd><kwd>K-means</kwd><kwd>computational efficiency</kwd></kwd-group><custom-meta-group><custom-meta><meta-name>File created by JATS Editor</meta-name><meta-value>https://jatseditor.com</meta-value></custom-meta><custom-meta><meta-name>issue-created-year</meta-name><meta-value>2026</meta-value></custom-meta></custom-meta-group></article-meta></front><body><sec id="sec-1"><title>1. INTRODUCTION</title><p>Advancements in mathematics, science, and technology has driven humans to digital and information era. This modern era leads to the exponential growth of creating, collecting, and utilizing data as stated in Kapadiya et al. <xref ref-type="bibr" rid="BIBR-1">[1]</xref> and Horvitz and Mitchell <xref ref-type="bibr" rid="BIBR-2">[2]</xref>. As data collection and processing have advanced, the need for efective methods to extract meaningful insights from this data has become increasingly apparent. The development of mathematical models that capable of handling complex dataset is a necessity as well. Hokanson <xref ref-type="bibr" rid="BIBR-3">[3]</xref> stated that among the various approaches to model data, nonlinear least squares problems have emerged as a critical area of focus, particularly for their applications in data fitting and parameter estimation across diverse fields such as engineering, economics, and machine learning.</p><p>Nonlinear Least Squares problems involve finding the best fit curve to a set of data points by minimizing the sum of the squares of the residuals, which usually called the objective function of this problem. This task for finding optimal model that has high accuracy or lower error and has computational eficiency led to the development of numerous of numerical methods [<xref ref-type="bibr" rid="BIBR-4">4</xref>, <xref ref-type="bibr" rid="BIBR-5">5</xref>]. In Nocedal <xref ref-type="bibr" rid="BIBR-6">[6]</xref> and Yuan <xref ref-type="bibr" rid="BIBR-5">[5]</xref>, Gradient-based numerical methods is one of those methods for solving nonlinear least squares problems.</p><p>With the growth of data created, recent advances have explored the integration of stochastic elements into these numerical methods to potentially enhance their performance and its applications [<xref ref-type="bibr" rid="BIBR-7">7</xref>, <xref ref-type="bibr" rid="BIBR-8">8</xref>, <xref ref-type="bibr" rid="BIBR-9">9</xref>, <xref ref-type="bibr" rid="BIBR-10">10</xref>]. One specific method, Stochastic Gradient Descent that introduce randomness into the optimization process can influence both accuracy and computational eficiency as in Robbins <xref ref-type="bibr" rid="BIBR-7">[7]</xref> and Kiefer <xref ref-type="bibr" rid="BIBR-8">[8]</xref>. The impact of such stochastic methods on gradient-based numerical methods remains a topic of ongoing research.</p><p>The Levenberg-Marquardt algorithm has been extensively studied for solving nonlinear least squares problems. Early applications include the analysis of EPR spectra (1996) <xref ref-type="bibr" rid="BIBR-11">[11]</xref> and the development of damping and undamping strategies (1997) <xref ref-type="bibr" rid="BIBR-12">[12]</xref>. More recent eforts have focused on improving the accuracy and computational eficiency of the algorithm. For example, Yan et al. proposed AdaLM (Adaptive Levenberg-Marquardt), which adapts the step size to enhance performance <xref ref-type="bibr" rid="BIBR-13">[13]</xref>. Building on this line of work, Hong et al. (2020) introduced a stochastic version of the Levenberg-Marquardt algorithm by using randomly selected subsets of data. <xref ref-type="bibr" rid="BIBR-14">[14]</xref>. In contrast, our approach introduces a new variation of stochastic sampling: at each iteration, a diferent random subset of the data is selected. Furthermore, this study incorporates K-Means clustering, using cluster centroids as representative data points, which diferentiates our method from previous work.</p><p>The main purpose of this research is to investigate the influence of stochastic in numerical methods that utilize gradients, especially for Levenberg-Marquardt method, on accuracy and computational eficiency. In this study, numerical result with data taken in clusters will be compared to Stochastic Levenberg-Marquardt.</p></sec><sec id="sec-2"><title>2. PROBLEM STATEMENT AND METHODS</title><sec id="sec-3"><title>2.1. Least Squares Problem.</title><p>When we have data <inline-formula><tex-math id="math-1"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( x _ { i } , y _ { i } ) , \ i = 1 , 2 , \ldots , m \end{document} ]]></tex-math></inline-formula>, the relationship of <inline-formula><tex-math id="math-2"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x _ { i } \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-3"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle y _ { i } \end{document} ]]></tex-math></inline-formula> can be written in equation (1)</p><disp-formula id="equation-1"><tex-math id="math-4"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle y _ {i} = M (\mathbf {p}, x _ {i}) + \epsilon_ {i},\tag{1} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-5"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } = [ p _ { 1 } , p _ { 2 } , \ldots , p _ { n } ] ^ { \mathrm { T } } \end{document} ]]></tex-math></inline-formula> is a vector consist of model parameters, <inline-formula><tex-math id="math-6"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle M ( \mathbf { p } , x _ { i } ) \end{document} ]]></tex-math></inline-formula> is a fitting model, and <inline-formula><tex-math id="math-7"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle e _ { i } \end{document} ]]></tex-math></inline-formula> is measurement error on the data ordinate.</p><p>For any choice of <inline-formula><tex-math id="math-8"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } \end{document} ]]></tex-math></inline-formula>, residuals can be computed by</p><disp-formula id="equation-2"><tex-math id="math-9"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle f _ {i} (\mathbf {p}) = y _ {i} - M (\mathbf {p}, x _ {i}).\tag{2} \end{document} ]]></tex-math></disp-formula><p>The least squares problem is to minimize the sum of the squared residuals by obtaining <inline-formula><tex-math id="math-10"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } ^ { * } \end{document} ]]></tex-math></inline-formula> as the minimizer. Madsen <xref ref-type="bibr" rid="BIBR-15">[15]</xref> define least squares problem in Definition <xref ref-type="custom" custom-type="reference-target" rid="anchor-bc7534c7-ecd0-4907-9e9b-df730598d3d6">2.1</xref><target id="anchor-bc7534c7-ecd0-4907-9e9b-df730598d3d6" target-type="reference-target"/></p><p><bold>Definition 2.1. </bold><italic>[Least Squares Problem] Find a local minimizer </italic><inline-formula><tex-math id="math-11"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } ^ { * } ~ f o r \end{document} ]]></tex-math></inline-formula></p><disp-formula id="equation-3"><tex-math id="math-12"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F (\mathbf {p}) = \frac {1}{2} \sum_ {i = 1} ^ {m} (f _ {i} (\mathbf {p})) ^ {2},\tag{3} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-13"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle f _ { i } : \mathbb { R } ^ { n } \mapsto \mathbb { R } , \ i = 1 , 2 , \ldots ,m \end{document} ]]></tex-math></inline-formula>  are given functions, and <inline-formula><tex-math id="math-14"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m \geq n \end{document} ]]></tex-math></inline-formula>.</p><p>As an objective function, <inline-formula><tex-math id="math-15"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F ( \mathbf { p } ) \end{document} ]]></tex-math></inline-formula> can also be called as loss function or cost function.</p></sec><sec id="sec-4"><title>2.2. Gradient Descent.</title><p>Gradient Descent is an iterative optimization numerical method that uses first order derivative approach with the aim of minimizing the objective function. In Ruder <xref ref-type="bibr" rid="BIBR-16">[16]</xref>, Gradient Descent is a way to minimize objective function <inline-formula><tex-math id="math-16"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F ( \mathbf { p } _ { k } ; \zeta ) \end{document} ]]></tex-math></inline-formula> parameterized by its model parameters <inline-formula><tex-math id="math-17"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } _ { k } \in \mathbb { R } ^ { n } \end{document} ]]></tex-math></inline-formula> by updating the value of param eters with opposite direction of gradient of objective function. In Tian et al. <xref ref-type="bibr" rid="BIBR-17">[17]</xref>, updating the value of parameters <inline-formula><tex-math id="math-18"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } \end{document} ]]></tex-math></inline-formula> using Gradient Descent can be written as in equation <xref ref-type="disp-formula" rid="equation-1">(4)</xref></p><disp-formula id="equation-4"><tex-math id="math-19"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {p} _ {k + 1} = \mathbf {p} - \boldsymbol {\tau} _ {k} \nabla F (\mathbf {p} _ {k}; \boldsymbol {\zeta}),\tag{4} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-20"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k \end{document} ]]></tex-math></inline-formula> denoted as iteration of the simulation. Symbol <inline-formula><tex-math id="math-21"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tau _ { k } \end{document} ]]></tex-math></inline-formula> called as learning rate which usually constant greater than zero that regulate updating steps using gradient. Symbol <inline-formula><tex-math id="math-22"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \zeta : ( x _ { i } , y _ { i } ) , \ i = 1 , 2 , . . . , m \end{document} ]]></tex-math></inline-formula> is all samples used when calculating <inline-formula><tex-math id="math-23"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \nabla F ( \mathbf { p } _ { k } ; \zeta ) \end{document} ]]></tex-math></inline-formula> for every iteration.</p><p>Garrigos and Gower <xref ref-type="bibr" rid="BIBR-18">[18]</xref> showed one way about convergence of Gradient Descent as written in Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-795fbde5-6bbb-45c3-8c1c-bb532b517977">2.3</xref><target id="anchor-13032e68-1e84-4eaf-b85b-5bc233405bab" target-type="reference-target"/></p><p><bold>Definition 2.2. </bold><italic>Let </italic><inline-formula><tex-math id="math-24"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F:\mathbb{R}^n\to\mathbb{R} \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-25"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L > 0. \end{document} ]]></tex-math></inline-formula><italic></italic><inline-formula><tex-math id="math-26"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \ F \end{document} ]]></tex-math></inline-formula><italic> is said </italic><inline-formula><tex-math id="math-27"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L −smooth \end{document} ]]></tex-math></inline-formula><italic> if it is differentiable and if  </italic><inline-formula><tex-math id="math-28"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \nabla F:\mathbb{R}^n\to \mathbb { R } ^ { n } \end{document} ]]></tex-math></inline-formula><italic> is </italic><inline-formula><tex-math id="math-29"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L - L i p s c h i t z . \end{document} ]]></tex-math></inline-formula></p><p>For all </p><disp-formula id="equation-5"><tex-math id="math-30"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x, y \in \mathbb {R} ^ {n}, \| \nabla F (x) - \nabla F (y) \| \leq L \| x - y \|. \end{document} ]]></tex-math></disp-formula><p><target id="anchor-795fbde5-6bbb-45c3-8c1c-bb532b517977" target-type="reference-target"/></p><p><bold>Theorem 2.3.</bold><italic>Suppose we want to minimize a diferentiable function </italic><inline-formula><tex-math id="math-31"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \end{document} ]]></tex-math></inline-formula><italic> and minimizer (argmin) </italic><inline-formula><tex-math id="math-32"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \neq \emptyset \end{document} ]]></tex-math></inline-formula><italic>. Assume </italic><inline-formula><tex-math id="math-33"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \end{document} ]]></tex-math></inline-formula><italic> convex and </italic><inline-formula><tex-math id="math-34"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L−smooth \end{document} ]]></tex-math></inline-formula><italic>, </italic><inline-formula><tex-math id="math-35"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L > 0 \end{document} ]]></tex-math></inline-formula><italic>. Let </italic><inline-formula><tex-math id="math-36"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( p ^ { k } ) _ { k \in \mathbb { N } } \end{document} ]]></tex-math></inline-formula><italic> be a sequence of iterates generated by Gradient Descent algorithm, with a learning rate satisfying </italic><inline-formula><tex-math id="math-37"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array} { r } { \mathrm { ~ ~ 0 ~ } < \tau _ { k } < \frac { \mathrm { ~ 1 ~ } } { L } } \end{array} \end{document} ]]></tex-math></inline-formula><italic>. Then, for all </italic><inline-formula><tex-math id="math-38"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle p ^ { \ast } \in \end{document} ]]></tex-math></inline-formula><italic> argmin </italic><inline-formula><tex-math id="math-39"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \end{document} ]]></tex-math></inline-formula><italic>, for all </italic><inline-formula><tex-math id="math-40"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k \in \mathbb N \end{document} ]]></tex-math></inline-formula><italic>, we have</italic></p><disp-formula id="equation-6"><tex-math id="math-41"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F (p ^ {k}) - \inf F \leq \frac {\left\| p ^ {0} - p ^ {*} \right\| ^ {2}}{2 \tau_ {k} k}. \end{document} ]]></tex-math></disp-formula><p>As the number of data samples increases, the associated computational cost also rises accordingly. Therefore, recent advances introduce randomness or stochastic elements into Gradient Descent. It is done by sampling the data uniformly. In Tian et al. <xref ref-type="bibr" rid="BIBR-17">[17]</xref> stated that Stochastic Gradient Descent computes gradient only uses one random sample in each iteration. Mini-Batch Stochastic Gradient Descent randomly sampled small-batch data <inline-formula><tex-math id="math-42"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal {B} _ {j} \end{document} ]]></tex-math></inline-formula> to compute gradient. These two variants save a lot of memory in computation which might lead to speed increased. In this study, Stochastic Gradient Descent term will be used interchangeably for those two variants. With <inline-formula><tex-math id="math-43"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal {B} _ {j} , ~ j = 1 , 2 , \ldots , m \end{document} ]]></tex-math></inline-formula> batch data, updating parameters value can be written by</p><disp-formula id="equation-7"><tex-math id="math-44"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {p} _ {k + 1} = \mathbf {p} - \boldsymbol {\tau} _ {k} \nabla F (\mathbf {p} _ {k}; \mathcal {B} _ {j}).\tag{5} \end{document} ]]></tex-math></disp-formula><p>To show one way about the convergence of Stochastic Gradient Descent, Garrigos and Gower <xref ref-type="bibr" rid="BIBR-18">[18]</xref> depicts the objective function as a sum of functions <inline-formula><tex-math id="math-45"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F(p)\overset{\mathrm{def}}{=} \frac{1}{d}\sum_{i=1}^{d}F_i(p) \end{document} ]]></tex-math></inline-formula>, where <inline-formula><tex-math id="math-46"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F_i:\mathbb{R}^n\to\mathbb{R} \end{document} ]]></tex-math></inline-formula> is bounded from below and argmin <inline-formula><tex-math id="math-47"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \neq \emptyset \end{document} ]]></tex-math></inline-formula>. They said sum of functions problem as Sum of Convex when <inline-formula><tex-math id="math-48"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F_i:\mathbb{R}^n\to\mathbb{R} \end{document} ]]></tex-math></inline-formula> is assumed to be convex. They said sum of functions problem as Sum of <inline-formula><tex-math id="math-49"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L _ { m a x } \end{document} ]]></tex-math></inline-formula>−smooth when <inline-formula><tex-math id="math-50"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F_i:\mathbb{R}^n\to\mathbb{R} \end{document} ]]></tex-math></inline-formula> is assumed to be <inline-formula><tex-math id="math-51"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L _ { i } { \mathrm { - s m o o t h } } \end{document} ]]></tex-math></inline-formula> and denoted <inline-formula><tex-math id="math-52"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L _ { m a x } \ { \stackrel { \mathrm { d e f } } { = } } \ \operatorname* { m a x } L _ { i } \end{document} ]]></tex-math></inline-formula>. Garrigos and Gower <xref ref-type="bibr" rid="BIBR-18">[18]</xref> showed one way about convergence of Stochastic Gradient Descent as written in Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-3065f771-fc18-42df-9392-e1bbfc4f3755">2.4</xref>.<target id="anchor-3065f771-fc18-42df-9392-e1bbfc4f3755" target-type="reference-target"/></p><p><bold>Theorem 2.4.</bold><italic>Let the problem hold Sum of </italic><inline-formula><tex-math id="math-53"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L _ { \mathrm { m a x } } { - s m o o t h } \end{document} ]]></tex-math></inline-formula><italic> and Sum of Convex assumptions. Consider </italic><inline-formula><tex-math id="math-54"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( p ^ { k } ) _ { k \in \mathbb { N } } \end{document} ]]></tex-math></inline-formula><italic> be a sequence of iterates generated by Stochastic Gradient Descent algorithm, with a learning rate satisfying </italic><inline-formula><tex-math id="math-55"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array} { r } { \mathrm { ~ 0 ~ < ~ } \tau _ { k } < \frac { 1 } { 2 \mathcal { L } _ { j } } } \end{array} \end{document} ]]></tex-math></inline-formula><italic>. It follows that</italic></p><disp-formula id="equation-8"><tex-math id="math-56"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbb {E} [ F (p ^ {k}) - \inf F ] \leq \frac {\left\| p ^ {0} - p ^ {*} \right\| ^ {2}}{2 \sum_ {t = 0} ^ {k - 1} \tau_ {t} (1 - 2 \tau_ {t} \mathcal {L} _ {j})} + \frac {\sum_ {t = 0} ^ {k - 1} \tau_ {t} ^ {2}}{\sum_ {t = 0} ^ {k - 1} \tau_ {t} (1 - 2 \tau_ {T} \mathcal {L} _ {j})} \sigma_ {j} ^ {*}, \end{document} ]]></tex-math></disp-formula><p>where </p><disp-formula id="equation-9"><tex-math id="math-57"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \bar {p} ^ {k} \stackrel {\mathrm{def}} {=} \sum_ {t = 0} ^ {k - 1} u _ {k, t} p ^ {k} \end{document} ]]></tex-math></disp-formula><p> with <inline-formula><tex-math id="math-58"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle u _ {k, t} \stackrel {\mathrm{def}} {=} \frac {\tau_ {t} (1 - 2 \tau_ {t} \mathcal {L} _ {j})}{\sum_ {i = 0} ^ {k - 1} \tau_ {i} (1 - 2 \tau_ {i} \mathcal {L} _ {j})}. \end{document} ]]></tex-math></inline-formula></p></sec><sec id="sec-5"><title>2.3. Levenberg-Marquardt.</title><p>Levenberg-Marquardt (LM) is an iterative optimization numerical method that uses second order derivative approach with the aim of minimizing the objective function. While Gradient Method solely uses first order derivative, this method leverage the second order derivative. As this method may be considered as an extension of Gauss-Newton’s method, the Levenberg-Marquardt also based on solving linearized least squares problems. In Madsen <xref ref-type="bibr" rid="BIBR-15">[15]</xref>, the Levenberg-Marquardt method is based on a linear approximation to the components of <inline-formula><tex-math id="math-59"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {f} \end{document} ]]></tex-math></inline-formula> in the neighborhood of <inline-formula><tex-math id="math-60"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {p} \end{document} ]]></tex-math></inline-formula>. For small <inline-formula><tex-math id="math-61"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ||\mathbf {h}|| \end{document} ]]></tex-math></inline-formula>, the Taylor expansion is</p><disp-formula id="equation-10"><tex-math id="math-62"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {f} (\mathbf {p} + \mathbf {h}) \approx \mathbf {l} (\mathbf {h}) \equiv \mathbf {f} (\mathbf {p}) + \mathbf {J} (\mathbf {p}) \mathbf {h}.\tag{6} \end{document} ]]></tex-math></disp-formula><p>Thus, <inline-formula><tex-math id="math-63"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F \end{document} ]]></tex-math></inline-formula> denoted as</p><disp-formula id="equation-11"><tex-math id="math-64"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{align*}F(\mathbf{p} + \mathbf{h}) \approx L(\mathbf{h}) &\equiv \frac{1}{2} \mathbf{l}(\mathbf{h})^{\mathrm{T}} \mathbf{l}(\mathbf{h}) \\&= \frac{1}{2} \mathbf{f}^{\mathrm{T}} \mathbf{f} + \mathbf{h}^{\mathrm{T}} \mathbf{J}^{\mathrm{T}} \mathbf{f} + \frac{1}{2} \mathbf{h}^{\mathrm{T}} \mathbf{J}^{\mathrm{T}} \mathbf{J} \mathbf{h} \tag{7} \\&= F(\mathbf{p}) + \mathbf{h}^{\mathrm{T}} \mathbf{J}^{\mathrm{T}} \mathbf{f} + \frac{1}{2} \mathbf{h}^{\mathrm{T}} \mathbf{J}^{\mathrm{T}} \mathbf{J} \mathbf{h},\end{align*} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-65"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { J } \end{document} ]]></tex-math></inline-formula> is Jacobian. Linear approximation for Gradient and Hessian respectively are</p><disp-formula id="equation-12"><tex-math id="math-66"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F ^ {\prime} (\mathbf {p}) = L ^ {\prime} (\mathbf {0}) = \mathbf {J} ^ {\mathrm{T}} \mathbf {f} \qquad F ^ {\prime \prime} (\mathbf {p}) = L ^ {\prime \prime} (\mathbf {0}) = \mathbf {J} ^ {\mathrm{T}} \mathbf {J}.\tag{8} \end{document} ]]></tex-math></disp-formula><p>Updating the value of <inline-formula><tex-math id="math-67"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } \end{document} ]]></tex-math></inline-formula> using <inline-formula><tex-math id="math-68"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { h } _ { l m } \end{document} ]]></tex-math></inline-formula> with Levenberg-Marquardt method is done by solving equation <xref ref-type="disp-formula" rid="equation-2">(9)</xref></p><disp-formula id="equation-13"><tex-math id="math-69"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle (\mathbf {J} ^ {\mathrm{T}} \mathbf {J} + \mu \mathbf {I}) \mathbf {h} _ {l m} = - \mathbf {g} \quad \text { with } \quad \mathbf {g} = \mathbf {J} ^ {\mathrm{T}} \mathbf {f} \text { and } \mu \geq 0.\tag{9} \end{document} ]]></tex-math></disp-formula><p>Constant <inline-formula><tex-math id="math-70"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mu \end{document} ]]></tex-math></inline-formula> is the damping parameter which influences both the direction and the size of the step <xref ref-type="bibr" rid="BIBR-15">[15]</xref>. It is stated that the initial value of <inline-formula><tex-math id="math-71"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mu \end{document} ]]></tex-math></inline-formula> is related to the size of the elements in <inline-formula><tex-math id="math-72"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle { \mathbf { A } _ { 0 } }  { = } \mathbf { J } ( \mathbf { p } _ { 0 } ) ^ { \mathrm { T } } \mathbf { J } ( \mathbf { p } _ { 0 } ) \end{document} ]]></tex-math></inline-formula> by letting</p><disp-formula id="equation-14"><tex-math id="math-73"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mu_ {0} = \theta \cdot \max _ {i} \{a _ {i i} ^ {(0)} \}\tag{10} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-74"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta \end{document} ]]></tex-math></inline-formula> is chosen by user. Updating value <inline-formula><tex-math id="math-75"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } \end{document} ]]></tex-math></inline-formula> is controlled by gain ratio <inline-formula><tex-math id="math-76"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle q \end{document} ]]></tex-math></inline-formula> in equation (11). If <inline-formula><tex-math id="math-77"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle q \end{document} ]]></tex-math></inline-formula> is large value, <inline-formula><tex-math id="math-78"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L ( \mathbf { h } _ { l m } ) \end{document} ]]></tex-math></inline-formula> is a good approximation on <inline-formula><tex-math id="math-79"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle F ( \mathbf { p } + \mathbf { h } _ { l m } ) \end{document} ]]></tex-math></inline-formula>. If <inline-formula><tex-math id="math-80"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle q \end{document} ]]></tex-math></inline-formula> is small, <inline-formula><tex-math id="math-81"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle L ( \mathbf { h } _ { l m } ) \end{document} ]]></tex-math></inline-formula> is a poor approximation.</p><disp-formula id="equation-15"><tex-math id="math-82"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle q = \frac {F (\mathbf {p}) - F (\mathbf {p} + \mathbf {h} _ {l m})}{L (\mathbf {0}) - L (\mathbf {h} _ {l m})}.\tag{11} \end{document} ]]></tex-math></disp-formula><p>The pseudocode which details the step for updating parameters value using Levenberg-Marquardt are shown in <xref ref-type="fig" rid="figure-1">Algorithm 1</xref>.</p><fig id="figure-1"><label>Algorithm 1</label><caption><p>Levenberg-Marquardt</p></caption><graphic xlink:href="https://www.jims-a.org/index.php/jimsa/article/download/1960/541/13862" mime-subtype="png" mimetype="image"><alt-text>Algorithm 1</alt-text></graphic></fig></sec><sec id="sec-6"><title>2.4. Stochastic Levenberg-Marquardt.</title><p>While Stochastic Gradient Descent randomly sampled data for calculating gradient (first order derivative), in Stochastic Levenberg-Marquardt (SLM), randomly sampled data used for calculating gradient as well as Hessian matrix, specifically for calculating Jacobian matrix (first order and second order derivative) <xref ref-type="bibr" rid="BIBR-19">[19]</xref>. In Levenberg-Marquardt, Jacobian matrix, gradient, and Hessian matrix are calculated with all data samples, while Stochastic Levenberg-Marquardt calculate those with randomly sampled data. The pseudocode that details the step for updating parameters value using Stochastic Levenberg-Marquardt are shown in <xref ref-type="fig" rid="figure-2">Algorithm 2</xref>.</p></sec><sec id="sec-7"><title>2.5. K-Means.</title><p>In this study, data will be sampled with centroid-based clustering algorithm, specifically K-Means. Centers from clustering algorithm will represent the data to get the optimized model. K-Means is preferred in this study since the number of centroids (K) could be specified. K-Means algorithm initially proposed by MacQueen <xref ref-type="bibr" rid="BIBR-20">[20]</xref> in 1967. This method classifies given object to <inline-formula><tex-math id="math-83"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k \end{document} ]]></tex-math></inline-formula> diferent clusters <xref ref-type="bibr" rid="BIBR-21">[21]</xref>.</p><p>K-Means algorithm used in this study is K-Means++. Let <inline-formula><tex-math id="math-84"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal { X } \end{document} ]]></tex-math></inline-formula> be data points, let <inline-formula><tex-math id="math-85"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle { \mathcal { C } } = \{ c _ { 1 } , c _ { 2 } , \ldots , c _ { k } \} \end{document} ]]></tex-math></inline-formula> are centers, and let <inline-formula><tex-math id="math-86"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle D ( x ) \end{document} ]]></tex-math></inline-formula> denote the shortest distance from a data point to the closest center which have already chosen, as stated in Arthur and Vassilvitskii <xref ref-type="bibr" rid="BIBR-22">[22]</xref> the steps for K-Means++ are</p><list list-type="order"><list-item><p>Take one center <inline-formula><tex-math id="math-87"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle c _ { 1 } \end{document} ]]></tex-math></inline-formula>, chosen uniformly at random from <inline-formula><tex-math id="math-88"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal { X } \end{document} ]]></tex-math></inline-formula>.</p></list-item><list-item><p>Take a new center <inline-formula><tex-math id="math-89"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle c _ { i } \end{document} ]]></tex-math></inline-formula>, choosing <inline-formula><tex-math id="math-90"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x \in \mathcal { X } \end{document} ]]></tex-math></inline-formula> with probability <inline-formula><tex-math id="math-91"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \frac { D ( x ) ^ { 2 } } { \sum _ { x \in \mathcal { X } } D ( x ) ^ { 2 } } \end{document} ]]></tex-math></inline-formula>.</p></list-item><list-item><p>Repeat step 2 until <inline-formula><tex-math id="math-92"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k \end{document} ]]></tex-math></inline-formula> centers have taken altogether.</p></list-item><list-item><p>For each <inline-formula><tex-math id="math-93"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle i \in \{ 1 , 2 , \ldots , k \} \end{document} ]]></tex-math></inline-formula>, set the cluster <inline-formula><tex-math id="math-94"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle C _ { i } \end{document} ]]></tex-math></inline-formula> to be the set of points in <inline-formula><tex-math id="math-95"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal { X } \end{document} ]]></tex-math></inline-formula> that are closer to <inline-formula><tex-math id="math-96"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle c _ { i } \end{document} ]]></tex-math></inline-formula> than they are to <inline-formula><tex-math id="math-97"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle c _ { j } \end{document} ]]></tex-math></inline-formula> for all <inline-formula><tex-math id="math-98"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle j \neq i \end{document} ]]></tex-math></inline-formula>.</p></list-item><list-item><p>For each <inline-formula><tex-math id="math-99"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle i \in \{ 1 , 2 , \ldots , k \} \end{document} ]]></tex-math></inline-formula>, set <inline-formula><tex-math id="math-100"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle c _ { i } \end{document} ]]></tex-math></inline-formula> to be the center of mass of all points in <inline-formula><tex-math id="math-101"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle C _ { i } \end{document} ]]></tex-math></inline-formula>.</p></list-item><list-item><p>Repeat step 4 and step 5 until there are no any changes in <inline-formula><tex-math id="math-102"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathcal { C } \end{document} ]]></tex-math></inline-formula>.</p></list-item></list><fig id="figure-2"><label>Algorithm 2</label><caption><p>Stochastic Levenberg-Marquardt</p></caption><graphic xlink:href="https://www.jims-a.org/index.php/jimsa/article/download/1960/541/13863" mime-subtype="png" mimetype="image"><alt-text>Algorithm 2</alt-text></graphic></fig></sec></sec><sec id="sec-8"><title>3. SIMULATION SETUP</title><sec id="sec-9"><title>3.1. Data and Preprocessing.</title><p>Apple’s Stock Price Data (AAPL) from December 12, 1980 to May 16, 2024 was used for building the model. The dataset was taken from <ext-link ext-link-type="uri" xlink:href="https://finance.yahoo.com/quote/AAPL/history/?period1=345479400&amp;period2=1715904000" xlink:title="Yahoo! Finance">Yahoo! Finance</ext-link> with 10,948 rows of data. In this study, <inline-formula><tex-math id="math-103"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 5 ^ { \mathrm { t h } } \end{document} ]]></tex-math></inline-formula> degree polynomial model <inline-formula><tex-math id="math-104"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l ( x ) = a + b x + cx^2+dx^3+ex^4+fx^5 \end{document} ]]></tex-math></inline-formula>, <inline-formula><tex-math id="math-105"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \quad x\in[0,10947] \end{document} ]]></tex-math></inline-formula> is built.</p><p>In the dataset, there was a diference in <inline-formula><tex-math id="math-106"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-107"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle y \end{document} ]]></tex-math></inline-formula> scale. With scale diferent, it caused the sensitivity for model’s parameters. Wan <xref ref-type="bibr" rid="BIBR-23">[23]</xref> stated that to obtain better convergence using Gradient Descent, the normalization method could help. Thus, in this study, the data were normalized with <italic>Min-Max Normalization</italic><xref ref-type="bibr" rid="BIBR-24">[24]</xref>. This method is chosen since the distribution of <inline-formula><tex-math id="math-108"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-109"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle y \end{document} ]]></tex-math></inline-formula> are not normal or close to normal. Min-Max Normalization is shown in equation <xref ref-type="disp-formula" rid="equation-3">(12)</xref></p><disp-formula id="equation-16"><tex-math id="math-110"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle x _ {\mathrm{new}} = \frac {x - \min (x)}{\max (x) - \min (x)}.\tag{12} \end{document} ]]></tex-math></disp-formula><p>Loss function used in this study is Mean Squared Error (MSE) which show in equation (13)</p><disp-formula id="equation-17"><tex-math id="math-111"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle M S E = \frac {1}{m} \sum_ {i = 1} ^ {m} f _ {i} (\mathbf {p}) ^ {2}.\tag{13} \end{document} ]]></tex-math></disp-formula><p>Apart from Levenberg-Marquardt simulation, Stochastic Levenberg-Marquardt with 1, 100, 200, 300, 400, and 500 randomly sampled data were simulated. For Cluster-Levenberg-Marquardt, cluster with 100, 200, 300, 400, and 500 centers were simulated. These numbers were selected after experimenting with random sample far below 100, which had similar result with 1 random sample, and above 500, which had similar result with 500 random samples.</p></sec><sec id="sec-10"><title>3.2. Parameters Setup.</title><p>Initial setups for Levenberg-Marquardt, Stochastic Levenberg-Marquardt and Cluster-Levenberg-Marquardt are</p><list list-type="order"><list-item><p>Maximum iteration = 50</p><p>In the next section, it will be shown that it does not require higher iteration since the fluctuation of loss function was similar in higher maximum iteration.</p></list-item><list-item><p><inline-formula><tex-math id="math-112"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta = 1 \times 1 0 ^ { - 7 } \end{document} ]]></tex-math></inline-formula></p><p>As a rule of thumb, small value of <inline-formula><tex-math id="math-113"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle θ \end{document} ]]></tex-math></inline-formula> is chosen, but the algorithm is not very sensitive by the choice of <inline-formula><tex-math id="math-114"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle θ \end{document} ]]></tex-math></inline-formula><xref ref-type="bibr" rid="BIBR-15">[15]</xref>.</p></list-item><list-item><p>Initial <inline-formula><tex-math id="math-115"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { p } = ( a , b , c , d , e , f ) = (−0.75276, 2.70429, 1.39196, 0.59195, −2.06389, −2.06403) \end{document} ]]></tex-math></inline-formula>.</p></list-item></list><p>The initial guess is uniform randomly generated. The initial guess could be chosen by the user or randomly generated.</p><p>Computer specifications for running Python simulation program was ASUS Desktop 12th Gen Intel(R) Core(TM) i7-12700, 2100 Mhz 12 Core(s) 20 Logical Processor(s).</p></sec></sec><sec id="sec-11"><title>4. SIMULATION RESULTS AND DISCUSSION</title><p>The loss function comparison of Levenberg-Marquardt, Stochastic Levenberg-Marquardt, and Cluster-Levenberg-Marquardt is displayed in <xref ref-type="fig" rid="figure-3">Figure 1</xref>. Levenberg-Marquardt has the lowest loss function since it computes the Jacobian using all sample data, as illustrated in <xref ref-type="fig" rid="figure-3">Figure 1</xref>. It is observed that after completing 50 iterations, the loss function of the stochastic Levenberg-Marquardt with 1 sample model does not approach Levenberg-Marquardt, as shown in <xref ref-type="fig" rid="figure-3">Figure 1</xref>a, where the loss function is significantly higher than Levenberg-Marquardt. The minimum MSE of SLM 1 sample from 50 iterations is significantly higher, as displayed in <xref ref-type="table" rid="table-1">Table 1</xref>, than the minimum MSE of LM. It is suspected that additional iterations are required for loss function of SLM 1 sample to get closer to LM.</p><table-wrap id="table-1"><label>Table 1</label><caption><p>Loss Function Comparison and Relative Error</p></caption><table><colgroup><col></col><col></col><col></col></colgroup><thead><tr><th scope="col">Method</th><th scope="col">Minimum MSE</th><th scope="col">Relative Error (%)</th></tr></thead><tbody><tr><td>LM</td><td>0.001574695911</td><td>0</td></tr><tr><td>SLM 1</td><td>0.239841356826</td><td>15130.96333412</td></tr><tr><td>SLM 100</td><td>0.001575429783</td><td>0.046604042483</td></tr><tr><td>SLM 200</td><td>0.001578011113</td><td>0.210529695116</td></tr><tr><td>SLM 300</td><td>0.001575371382</td><td>0.042895308592</td></tr><tr><td>SLM 400</td><td>0.001576264816</td><td>0.099632261320</td></tr><tr><td>SLM 500</td><td>0.001575898006</td><td>0.076338215123</td></tr><tr><td>Cl 100</td><td>0.001610635152</td><td>2.282297240011</td></tr><tr><td>Cl 200</td><td>0.001606480400</td><td>2.018452474183</td></tr><tr><td>Cl 300</td><td>0.001617227325</td><td>2.700928708518</td></tr><tr><td>Cl 400</td><td>0.001613599200</td><td>2.470527099832</td></tr><tr><td>Cl 500</td><td>0.001612992396</td><td>2.431992407551</td></tr></tbody></table></table-wrap><p>Next, <xref ref-type="fig" rid="figure-3">Figure 1</xref>b, <xref ref-type="fig" rid="figure-3">Figure 1</xref>c, <xref ref-type="fig" rid="figure-3">Figure 1</xref>d, <xref ref-type="fig" rid="figure-3">Figure 1</xref>e, and <xref ref-type="fig" rid="figure-3">Figure 1</xref>f compares loss function for Levenberg-Marquardt, Stochastic Levenberg-Marquardt with 100 samples, and Cluster-Levenberg-Marquardt with 100 centers, and so forth. With Cluster-Levenberg-Marquardt are slightly higher than Levenberg-Marquardt. However, it never approach Levenberg-Marquardt loss function. In <xref ref-type="table" rid="table-1">Table 1</xref>, the relative error of the minimum MSE of Cluster Levenberg-Marquardt to the minimum MSE of Levenberg-Marquardt are below 2.71%. On the other hand, because of the stochastic elements in iterations, in Stochastic Levenberg-Marquardt, there are some iterations has MSE are high, even higher than Cluster-Levenberg-Marquardt. Nevertheless, in certain instances, the MSE of Stochastic Levenberg-Marquardt approaches the MSE of Levenberg-Marquardt. In <xref ref-type="table" rid="table-1">Table 1</xref>, the relative error of the minimum MSE of Stochastic Levenberg-Marquardt to the minimum MSE of Levenberg-Marquardt are below 0.22%. Thus, the accuracy of Stochastic Levenberg-Marquardt is superior than the Cluster Levenberg-Marquardt.</p><p><xref ref-type="fig" rid="figure-4">Figure 2</xref> shows the 5<sup>th</sup> degree polynomial curve with Levenberg-Marquardt, Stochastic Levenberg-Marquardt, and Cluster Levenberg-Marquardt. It is dificult to see any diferences in the curve, except SLM 1 sample, as the relative error are below 2.71%.</p><p><xref ref-type="fig" rid="figure-5">Figure 3</xref> shows the computation time comparison among the variants of methods and <xref ref-type="table" rid="table-2">Table 2</xref> shows total iteration of optimization simulation. For Cluster Levenberg-Marquardt, merely cluster with 100 centers have a lower cumulative computation time compared to Levenberg-Marquardt. In contrast, clusters with 200, 300, 400, and 500 centers have a higher cumulative computation time since K-Means clustering requires a large amount of computation time. On the other hand, despite the fact that the total number of iterations required for the simulation reaches the maximum number, Stochastic Levenberg-Marquardt has lower computation time compared to Levenberg-Marquardt.</p><fig id="figure-3"><label>Figure 1.</label><caption><p>Loss Comparison of Levenberg-Marquardt, Stochastic Levenberg-Marquardt, and Cluster-Levenberg-Marquardt</p></caption><graphic xlink:href="https://www.jims-a.org/index.php/jimsa/article/download/1960/541/13864" mime-subtype="png" mimetype="image"><alt-text>Figure 1.</alt-text></graphic></fig><fig id="figure-4"><label>Figure 2.</label><caption><p>Curve of 5th Degree Polynomial Model Comparison</p></caption><graphic xlink:href="https://www.jims-a.org/index.php/jimsa/article/download/1960/541/13865" mime-subtype="png" mimetype="image"><alt-text>Figure 2.</alt-text></graphic></fig><table-wrap id="table-2"><label>Table 2</label><caption><p>Iteration Comparison</p></caption><table><colgroup><col></col><col></col></colgroup><thead><tr><th scope="col">Method</th><th scope="col">Total Iteration</th></tr></thead><tbody><tr><td>LM</td><td>24</td></tr><tr><td>SLM 1</td><td>50</td></tr><tr><td>SLM 100</td><td>50</td></tr><tr><td>SLM 200</td><td>50</td></tr><tr><td>SLM 300</td><td>50</td></tr><tr><td>SLM 400</td><td>50</td></tr><tr><td>SLM 500</td><td>50</td></tr><tr><td>Cl 100</td><td>22</td></tr><tr><td>Cl 200</td><td>25</td></tr><tr><td>Cl 300</td><td>24</td></tr><tr><td>Cl 400</td><td>23</td></tr><tr><td>Cl 500</td><td>24</td></tr></tbody></table></table-wrap><p>In general, the computation time of Stochastic Levenberg-Marquardt and Cluster-Levenberg-Marquardt (excluding K-Means computation) are lower than Levenberg-Marquardt as shown in <xref ref-type="fig" rid="figure-5">Figure 3</xref>. Reducing computational time comes at the cost of accuracy, as these methods typically use smaller sample sizes. Smaller sample sizes contributes in faster computational cost for calculating Jacobian matrix, gradient, and Hessian matrix. Although computational time is reduced, the relative error compared to the Levenberg-Marquardt method remains low as the relative error are below 0.22% for the Stochastic Levenberg-Marquardt approach and below 2.71% for Cluster-Levenberg-Marquardt. With a small relative error compared to the Levenberg-Marquardt method and reduced computation time, this research demonstrates that Stochastic Levenberg-Marquardt outperforms to Levenberg-Marquardt and Cluster-Levenberg-Marquardt in computational eficiency.</p><fig id="figure-5"><label>Figure 3.</label><caption><p>Computation Time Comparison</p></caption><graphic xlink:href="https://www.jims-a.org/index.php/jimsa/article/download/1960/541/13866" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 3.</alt-text></graphic></fig></sec><sec id="sec-12"><title>5. CONCLUSION</title><p>This research aims to investigate the influence of stochastic in the Levenberg-Marquardt numerical method and sampling using cluster centers on accuracy and computational eficiency. Stochastic Levenberg-Marquardt with 100, 200, 300, 400, and 500 have relative error to Levenberg-Marquardt below 0.22% and have lower computational time in 50 maximum iteration. Cluster Levenberg-Marquardt with 100, 200, 300, 400, and 500 centers have relative error to Levenberg-Marquardt below 2.71% and have higher computational time, except 100 centers. To summarize, Stochastic Levenberg-Marquardt with 100, 200, 300, 400, and 500 outperformed Levenberg-Marquardt and Cluster-Levenberg-Marquardt with very low relative error and reduced computational time, which suggests that using Stochastic Levenberg-Marquardt approach is advantageous. Finally, further research may include a better stopping criterion for Stochastic Levenberg-Marquardt to stop before maximum iteration, applying the methods in bigger datasets, and exploring other clustering methods for better results in Levenberg-Marquardt.</p></sec></body><back><ack><title>Acknowledgement</title><p>This work is supported by Department of Mathematics, Faculty of Mathematics and Natural Sciences, Institut Teknologi Bandung.</p></ack><ref-list><title>REFERENCES</title><ref id="BIBR-1"><element-citation publication-type="conf-paper"><article-title>A study on largescale applications of big data in modern era</article-title><source>Proceedings of the 5th International Conference on Information Management &amp; Machine Intelligence, ACM</source><person-group person-group-type="author"><name><surname>Kapadiya</surname><given-names>D.</given-names></name><name><surname>Shekhawat</surname><given-names>C.</given-names></name><name><surname>Sharma</surname><given-names>P.</given-names></name></person-group><year>2023</year><pub-id pub-id-type="doi">10.1145/3647444.3647880</pub-id></element-citation></ref><ref id="BIBR-2"><element-citation publication-type="webpage"><article-title>From data to knowledge to action: A global enabler for the 21st century</article-title><person-group person-group-type="author"><name><surname>Horvitz</surname><given-names>E.</given-names></name><name><surname>Mitchell</surname><given-names>T.</given-names></name></person-group><year>2020</year><pub-id pub-id-type="doi">10.48550/ARXIV.2008.00045</pub-id></element-citation></ref><ref id="BIBR-3"><element-citation publication-type="journal"><article-title>Projected nonlinear least squares for exponential fitting</article-title><source>SIAM Journal on Scientific Computing</source><volume>39</volume><issue>6</issue><person-group person-group-type="author"><name><surname>Hokanson</surname><given-names>J.M.</given-names></name></person-group><year>2017</year><page-range>3107-3128,</page-range><pub-id pub-id-type="doi">10.1137/16M1084067</pub-id></element-citation></ref><ref id="BIBR-4"><element-citation publication-type="journal"><article-title>Least squares optimization: From theory to practice</article-title><source>Robotics</source><volume>9</volume><issue>3</issue><person-group person-group-type="author"><name><surname>Grisetti</surname><given-names>G.</given-names></name><name><surname>Guadagnino</surname><given-names>T.</given-names></name><name><surname>Aloise</surname><given-names>I.</given-names></name><name><surname>Colosi</surname><given-names>M.</given-names></name><name><surname>Corte</surname><given-names>B.Della</given-names></name><name><surname>Schlegel</surname><given-names>D.</given-names></name></person-group><year>2020</year><page-range>51,</page-range><pub-id pub-id-type="doi">10.3390/robotics9030051</pub-id></element-citation></ref><ref id="BIBR-5"><element-citation publication-type="journal"><article-title>Recent advances in numerical methods for nonlinear equations and nonlinear least squares</article-title><source>Numerical Algebra, Control &amp; Optimization</source><volume>1</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Yuan</surname><given-names>Y.-X.</given-names></name></person-group><year>2011</year><page-range>15-34,</page-range><pub-id pub-id-type="doi">10.3934/naco.2011.1.15</pub-id></element-citation></ref><ref id="BIBR-6"><element-citation publication-type="book"><article-title>Numerical Optimization</article-title><person-group person-group-type="author"><name><surname>Nocedal</surname><given-names>J.</given-names></name><name><surname>Wright</surname><given-names>S.J.</given-names></name></person-group><year>2006</year><publisher-name>Springer</publisher-name><pub-id pub-id-type="doi">10.1007/978-0-387-40065-5</pub-id></element-citation></ref><ref id="BIBR-7"><element-citation publication-type="journal"><article-title>A stochastic approximation method</article-title><source>The Annals of Mathematical Statistics</source><volume>22</volume><issue>3</issue><person-group person-group-type="author"><name><surname>Robbins</surname><given-names>H.</given-names></name><name><surname>Monro</surname><given-names>S.</given-names></name></person-group><year>1951</year><page-range>400-407,</page-range><pub-id pub-id-type="doi">10.1214/aoms/1177729586</pub-id></element-citation></ref><ref id="BIBR-8"><element-citation publication-type="journal"><article-title>Stochastic estimation of the maximum of a regression function</article-title><source>The Annals of Mathematical Statistics</source><volume>23</volume><issue>3</issue><person-group person-group-type="author"><name><surname>Kiefer</surname><given-names>J.</given-names></name><name><surname>Wolfowitz</surname><given-names>J.</given-names></name></person-group><year>1952</year><page-range>462-466,</page-range><pub-id pub-id-type="doi">10.1214/aoms/1177729392</pub-id></element-citation></ref><ref id="BIBR-9"><element-citation publication-type="journal"><article-title>A stochastic levenberg– marquardt method using random models with complexity results</article-title><source>SIAM/ASA Journal on Uncertainty Quantification</source><volume>10</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Bergou</surname><given-names>E.H.</given-names></name><name><surname>Diouane</surname><given-names>Y.</given-names></name><name><surname>Kungurtsev</surname><given-names>V.</given-names></name><name><surname>Royer</surname><given-names>C.W.</given-names></name></person-group><year>2022</year><page-range>507-536,</page-range><pub-id pub-id-type="doi">10.1137/20M1366253</pub-id></element-citation></ref><ref id="BIBR-10"><element-citation publication-type="journal"><article-title>Stochastic levenberg–marquardt neural network implementation for analyzing the convective heat transfer in a wavy fin</article-title><source>Mathematics</source><volume>11</volume><issue>10</issue><person-group person-group-type="author"><name><surname>Kumar</surname><given-names>R.S.V.</given-names></name><name><surname>Alsulami</surname><given-names>M.D.</given-names></name><name><surname>Sarris</surname><given-names>I.E.</given-names></name><name><surname>Sowmya</surname><given-names>G.</given-names></name><name><surname>Gamaoun</surname><given-names>F.</given-names></name></person-group><year>2023</year><page-range>2401,</page-range><pub-id pub-id-type="doi">10.3390/math11102401</pub-id></element-citation></ref><ref id="BIBR-11"><element-citation publication-type="journal"><article-title>Nonlinear-least-squares analysis of slowmotion epr spectra in one and two dimensions using a modified levenberg–marquardt algorithm</article-title><source>Journal of Magnetic Resonance Series A</source><volume>120</volume><issue>2</issue><person-group person-group-type="author"><name><surname>Budil</surname><given-names>D.E.</given-names></name><name><surname>Lee</surname><given-names>S.</given-names></name><name><surname>Saxena</surname><given-names>S.</given-names></name><name><surname>Freed</surname><given-names>J.H.</given-names></name></person-group><year>1996</year><page-range>155-189,</page-range><pub-id pub-id-type="doi">10.1006/jmra.1996.0113</pub-id></element-citation></ref><ref id="BIBR-12"><element-citation publication-type="journal"><article-title>Damping–undamping strategies for the levenberg–marquardt nonlinear leastsquares method</article-title><source>Computers in Physics</source><volume>11</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Lampton</surname><given-names>M.</given-names></name></person-group><year>1997</year><page-range>110-115,</page-range><pub-id pub-id-type="doi">10.1063/1</pub-id></element-citation></ref><ref id="BIBR-13"><element-citation publication-type="journal"><article-title>Adaptive levenberg–marquardt algorithm: A new optimization strategy for levenberg–marquardt neural networks</article-title><source>Mathematics</source><volume>9</volume><issue>17</issue><person-group person-group-type="author"><name><surname>Yan</surname><given-names>Z.</given-names></name><name><surname>Zhong</surname><given-names>S.</given-names></name><name><surname>Lin</surname><given-names>L.</given-names></name><name><surname>Cui</surname><given-names>Z.</given-names></name></person-group><year>2021</year><page-range>2176,</page-range><pub-id pub-id-type="doi">10.3390/math9172176</pub-id></element-citation></ref><ref id="BIBR-14"><element-citation publication-type="conf-paper"><article-title>Stochastic levenberg–marquardt for solving optimization problems on hardware accelerators</article-title><source>IEEE Conference Proceedings, IEEE</source><person-group person-group-type="author"><name><surname>Hong</surname><given-names>Y.</given-names></name><etal/></person-group><year>2020</year><ext-link xlink:href="https://repository" ext-link-type="uri" xlink:title="Website link">Website link</ext-link></element-citation></ref><ref id="BIBR-15"><element-citation publication-type="journal"><article-title>Methods for Non-Linear Least Squares Problems</article-title><person-group person-group-type="author"><name><surname>Madsen</surname><given-names>K.</given-names></name><name><surname>Nielsen</surname><given-names>H.B.</given-names></name><name><surname>Tinglef</surname><given-names>O.</given-names></name></person-group><year>2004</year><edition>2</edition></element-citation></ref><ref id="BIBR-16"><element-citation publication-type="journal"><article-title>An overview of gradient descent optimization algorithms</article-title><person-group person-group-type="author"><name><surname>Ruder</surname><given-names>S.</given-names></name></person-group><year>2016</year><comment>// org/10.48550/ARXIV.1609.04747</comment></element-citation></ref><ref id="BIBR-17"><element-citation publication-type="journal"><article-title>Recent advances in stochastic gradient descent in deep learning</article-title><source>Mathematics</source><volume>11</volume><issue>3</issue><person-group person-group-type="author"><name><surname>Tian</surname><given-names>Y.</given-names></name><name><surname>Zhang</surname><given-names>Y.</given-names></name><name><surname>Zhang</surname><given-names>H.</given-names></name></person-group><year>2023</year><page-range>682,</page-range><pub-id pub-id-type="doi">10.3390/math11030682</pub-id></element-citation></ref><ref id="BIBR-18"><element-citation publication-type="webpage"><article-title>Handbook of convergence theorems for (stochastic) gradient methods</article-title><person-group person-group-type="author"><name><surname>Garrigos</surname><given-names>G.</given-names></name><name><surname>Gower</surname><given-names>R.M.</given-names></name></person-group><year>2023</year><pub-id pub-id-type="doi">10.48550/ARXIV.2301.11235</pub-id></element-citation></ref><ref id="BIBR-19"><element-citation publication-type="journal"><article-title>Global convergence of a stochastic levenberg–marquardt algorithm based on trust region</article-title><source>Journal of the Operations Research Society of China</source><person-group person-group-type="author"><name><surname>Shao</surname><given-names>W.-Y.</given-names></name><name><surname>Fan</surname><given-names>J.-Y.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.1007/s40305-023-00529-6</pub-id></element-citation></ref><ref id="BIBR-20"><element-citation publication-type="conf-paper"><article-title>Some methods for classification and analysis of multivariate observations</article-title><source>Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability</source><volume>1</volume><person-group person-group-type="author"><name><surname>MacQueen</surname><given-names>J.B.</given-names></name></person-group><year>1967</year><page-range>281-297,</page-range><ext-link xlink:href="https://digitalassets.lib.berkeley.edu/math/ucb/text/math" ext-link-type="uri" xlink:title="Math">Math</ext-link></element-citation></ref><ref id="BIBR-21"><element-citation publication-type="conf-paper"><article-title>Research on k-means clustering algorithm: An improved k-means clustering algorithm</article-title><source>2010 Third International Symposium on Intelligent Information Technology and Security Informatics</source><person-group person-group-type="author"><name><surname>Na</surname><given-names>S.</given-names></name><name><surname>Xumin</surname><given-names>L.</given-names></name><name><surname>Yong</surname><given-names>G.</given-names></name></person-group><year>2010</year><page-range>63-67,</page-range><publisher-name>IEEE</publisher-name><pub-id pub-id-type="doi">10.1109/IITSI.2010.74</pub-id></element-citation></ref><ref id="BIBR-22"><element-citation publication-type="conf-paper"><article-title>k-means++: The advantages of careful seeding</article-title><source>Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms</source><person-group person-group-type="author"><name><surname>Arthur</surname><given-names>D.</given-names></name><name><surname>Vassilvitskii</surname><given-names>S.</given-names></name></person-group><year>2007</year><page-range>1027-1035,</page-range><pub-id pub-id-type="doi">10.5555/1283383.1283494</pub-id></element-citation></ref><ref id="BIBR-23"><element-citation publication-type="journal"><article-title>Influence of feature scaling on convergence of gradient iterative algorithm</article-title><source>Journal of Physics: Conference Series</source><volume>1213</volume><issue>3</issue><person-group person-group-type="author"><name><surname>Wan</surname><given-names>X.</given-names></name></person-group><year>2019</year><page-range>032021,</page-range><pub-id pub-id-type="doi">10.1088/1742696/1213/3/032021</pub-id></element-citation></ref><ref id="BIBR-24"><element-citation publication-type="journal"><article-title>Impact of data normalization on stock index forecasting</article-title><source>International Journal of Computer Information Systems and Industrial Management Applications</source><volume>6</volume><person-group person-group-type="author"><name><surname>Nayak</surname><given-names>S.C.</given-names></name><name><surname>Misra</surname><given-names>B.B.</given-names></name><name><surname>Behera</surname><given-names>H.S.</given-names></name></person-group><year>2014</year><page-range>357-369,</page-range></element-citation></ref></ref-list></back></article>