Best Machine Learning Libraries for Java in 2026

Compare ten machine learning libraries for Java by version, license, minimum Java and maintenance status, with Smile and Tribuo examples and a pick-by-situation table.

Decision map for Java machine learning libraries. If the model is trained in Python, use ONNX Runtime for ONNX files, DJL for PyTorch or Hugging Face models and TensorFlow Java for SavedModels. If the model is trained in Java on data that fits in memory, use Smile when GPL or a commercial license is fine and Tribuo when an Apache license is required. If the data needs a cluster, use Spark MLlib. For learning, use Weka.

For classic machine learning on tabular data in Java, the strongest libraries in 2026 are Smile and Tribuo, and for running deep learning models that were trained in Python, the standard choices are DJL and ONNX Runtime. Weka still fits teaching and quick experiments, and Spark MLlib fits data that does not fit on one machine.

A Java team picks a machine learning library for one of two jobs. Either we train a model inside the JVM on our own data, or a data science team trains the model in Python and we serve it from a Java service. The two jobs need different libraries, and many top-10 lists mix them together without saying so.

Training a model in Java can take a few lines. The next snippet trains a logistic regression classifier with Smile on an array of features and labels.

// Features: weight in grams, length in cm. Labels: 0 = apple, 1 = banana
double[][] x = {
    {150, 7}, {170, 8}, {140, 7}, {160, 8},
    {120, 18}, {130, 20}, {110, 17}, {125, 19}};
int[] y = {0, 0, 0, 0, 1, 1, 1, 1};

LogisticRegression model = LogisticRegression.fit(x, y);
int fruit1 = model.predict(new double[] {155, 7});    // 0, an apple

In the rest of the article, we compare ten Java machine learning libraries by the latest version, the license, the minimum Java version and whether the project still ships releases. After that, we train the same classifier with Smile and Tribuo, and we finish with a table that maps common situations to a library. Readers who already know their situation can jump to that table in section 7.

1. Java Machine Learning Libraries Compared

We checked every version and release date against Maven Central in October 2026. A library counts as active when it had a release in the last 12 months, because a stalled project stops receiving fixes for new Java versions and native platforms.

LibraryJobBest forLatest version (date)LicenseMin JavaStatus
SmileTrainClassic ML, statistics and NLP6.3.0 (Aug 2026)GPLv3 or commercial25Active
TribuoTrainClassic ML with tracked model provenance4.3.2 (Apr 2025)Apache-2.08 (17 for some modules)No release for 18 months
WekaTrainTeaching, GUI experiments, classic ML3.8.7 (Jul 2026)GPL-3.0Not statedActive
DJLInferenceDeep learning inference with PyTorch, TensorFlow or ONNX engines0.38.0 (Sep 2026)Apache-2.011Active
ONNX RuntimeInferenceFast inference for ONNX models1.31.0 (Oct 2026)MIT8Active
TensorFlow JavaInferenceRunning TensorFlow SavedModels1.2.0 (Jul 2026)Apache-2.011Active
Spark MLlibTrain on a clusterDistributed training on a cluster4.2.0 (Jul 2026)Apache-2.017, 21 or 25Active
Deeplearning4jAvoidDeep learning on the JVM1.0.0-M2.1 (Aug 2022)Apache-2.0Not statedNo release for 4 years
JlamaAvoidRunning LLMs in pure Java0.8.4 (Jan 2025)Apache-2.021 (preview features)No release for 21 months
MOAAvoidStream mining and concept drift2024.07.0 (Jul 2024)GPL-3.08No release for 27 months

Three libraries in the table use GPL-3.0, and Smile also sells a commercial license, so we check the license before we add any of them to closed-source software. The Apache-2.0 and MIT libraries have no such condition.

2. Libraries for Training Classic Models

Classic machine learning means algorithms such as logistic regression, random forests, gradient boosting and k-means on rows of numbers or categories. These models train in seconds on a laptop and are often the right first model for fraud scores or churn prediction.

2.1. Smile

Smile (Statistical Machine Intelligence and Learning Engine) has one of the largest sets of algorithms in Java, from classification and regression to clustering, manifold learning, NLP and plotting. It is fast, and its API takes plain Java arrays, so we can start without learning a data model first.

FactDetails
Maven coordinatescom.github.haifengl:smile-core
Latest version6.3.0, released August 2026
Minimum Java25 for Smile 5 and later (Smile 4 needs 21)
LicenseGPLv3 or a commercial license from HALO.AI
StrengthsA large set of algorithms, good speed, a Scala and Kotlin API, a DataFrame class and a shell
LimitsThe license rules out closed-source use without a paid license, and the Java 25 baseline rules out older runtimes

The Smile license file asks for a commercial license from any closed-source use that does not release its source under GPLv3, and it names internal enterprise apps and SaaS products as examples. The Maven POM only says GNU 3, so a dependency scanner may not flag the commercial terms.

Pick Smile if the project is open source or the company buys a license, and the team wants many algorithms behind one simple API.

2.2. Tribuo

Tribuo is a machine learning library from Oracle Labs. Every Tribuo model records its provenance, which is the dataset and the trainer with all hyperparameters used to build it, so we can audit a model in production or retrain it the same way a year later.

FactDetails
Maven coordinatesorg.tribuo:tribuo-all, or single modules such as tribuo-classification-sgd
Latest version4.3.2, released April 2025 (the main branch is on 5.0.0-SNAPSHOT)
Minimum Java8 for the core, 17 for the model card modules
LicenseApache-2.0
StrengthsTyped outputs (Label, Regressor), provenance, wrappers for XGBoost, ONNX Runtime and TensorFlow
LimitsNo release in 18 months and fewer algorithms than Smile

Pick Tribuo if the model must be auditable or the legal team needs an Apache license, and the team accepts the slower release pace.

2.3. Weka

Weka from the University of Waikato is the oldest library in the list, and the Data Mining textbook by Witten, Frank, Hall and Pal is built around it. Its Explorer GUI loads a CSV or ARFF file and compares algorithms without any code, and the same algorithms are available as a Java API.

FactDetails
Maven coordinatesnz.ac.waikato.cms.weka:weka-stable
Latest version3.8.7, released July 2026 (the development line is 3.9.7)
Minimum JavaNot stated on the download page, and the installers bundle OpenJDK 25
LicenseGPL-3.0
StrengthsA GUI for experiments, plus decades of examples and books
LimitsGPL license, and an older API built around the Instances class

Some articles link Weka to weka.io, which is an unrelated storage company. The library lives at the University of Waikato site linked above.

Pick Weka if we are learning machine learning or want to compare algorithms in a GUI before writing code.

3. Libraries for Running Deep Learning Models

Most deep learning models are trained in Python with PyTorch or TensorFlow. The Java side loads the trained file and runs inference, so these libraries are thin layers over a native engine written in C++. Each one downloads platform-specific native binaries, which we must account for in Docker images and on ARM servers.

3.1. Deep Java Library (DJL)

DJL is an engine-agnostic deep learning library from AWS. We write the Java code once against DJL’s API, and DJL runs the model on PyTorch, TensorFlow or ONNX Runtime underneath, which makes it the most flexible choice for serving models in Java.

FactDetails
Maven coordinatesai.djl:api plus an engine such as ai.djl.pytorch:pytorch-engine
Latest version0.38.0, released September 2026
Minimum Java11
LicenseApache-2.0
StrengthsSeveral engines behind one API, a model zoo, Hugging Face tokenizers, DJL Serving for model servers
LimitsStill on 0.x versions, and the native engines add hundreds of megabytes

Pick DJL if we serve PyTorch or Hugging Face models from Java and want the option to switch engines later.

3.2. ONNX Runtime for Java

ONNX Runtime from Microsoft runs models saved in the ONNX format, which PyTorch, TensorFlow, scikit-learn and XGBoost can all export. The Java API is small, with an OrtEnvironment, an OrtSession per model and tensors for inputs and outputs.

FactDetails
Maven coordinatescom.microsoft.onnxruntime:onnxruntime, or onnxruntime_gpu for CUDA
Latest version1.31.0, released October 2026
Minimum Java8
LicenseMIT
StrengthsFast CPU inference, GPU support, one format for models from many Python frameworks
LimitsInference only, and we write the pre- and post-processing ourselves

Pick ONNX Runtime if the data science team can export to ONNX and we want the smallest dependency for fast inference.

3.3. TensorFlow Java

TensorFlow Java is the official Java binding for TensorFlow, maintained by the SIG JVM group. It loads a TensorFlow SavedModel and can also train, though few teams train in Java with it.

FactDetails
Maven coordinatesorg.tensorflow:tensorflow-core-platform
Latest version1.2.0, released July 2026
Minimum Java11
LicenseApache-2.0
StrengthsRuns SavedModels without conversion, with the full TensorFlow op set
LimitsThe last Windows x86_64 build is 1.1.0, macOS x86_64 builds ended before 1.1.0, and the README lags behind the releases

Pick TensorFlow Java if the models are TensorFlow SavedModels and converting them to ONNX is not an option.

4. Spark MLlib for Data That Needs a Cluster

Spark MLlib trains models on DataFrames spread across a Spark cluster. It has the classic algorithms, along with pipelines for feature engineering, and it runs the same code on a laptop or on hundreds of nodes. The DataFrame API in the spark.ml package is the one to use, because the older RDD API is in maintenance mode.

FactDetails
Maven coordinatesorg.apache.spark:spark-mllib_2.13
Latest version4.2.0, released July 2026
Minimum Java17, and Spark 4.2 also supports 21 and 25
LicenseApache-2.0
StrengthsScales to billions of rows, pipelines, model persistence, a Java API
LimitsLocal mode works for development, but Spark only pays off on a cluster, and its overhead is not worth it for data that fits in memory

Pick Spark MLlib if the training data is already in a data lake or does not fit on one machine.

5. Libraries to Avoid for New Projects

Older lists still recommend libraries that have not had a release for years. They work with the Java versions they were built for, but nobody fixes their bugs or security issues, and they fall behind new JDKs and CPU architectures.

LibraryLast releaseUse instead
Deeplearning4j1.0.0-M2.1, August 2022DJL or ONNX Runtime for inference
Jlama0.8.4, January 2025A hosted LLM through Spring AI or LangChain4j, or a local model server such as Ollama
MOA2024.07.0, July 2024Retraining a Smile or Tribuo model on a schedule. For true online learning, MOA still works, but it gets no fixes

Deeplearning4j’s README still says the project is in active development, but no build has reached Maven Central since 2022. If a new release appears, it is worth a second look.

6. Training the Same Classifier with Smile and Tribuo

The following example is a fruit classifier that tells apples from bananas by weight in grams and length in centimeters. We build it twice, with Smile 6.3.0 and Tribuo 4.3.2 on Java 25, because the two libraries show the trade-off between a short API and a typed one.

<dependency>
  <groupId>com.github.haifengl</groupId>
  <artifactId>smile-core</artifactId>
  <version>6.3.0</version>
</dependency>
<dependency>
  <groupId>org.tribuo</groupId>
  <artifactId>tribuo-classification-sgd</artifactId>
  <version>4.3.2</version>
</dependency>
<dependency>
  <groupId>org.slf4j</groupId>
  <artifactId>slf4j-nop</artifactId>
  <version>2.0.17</version>
</dependency>

Smile 6 is compiled for Java 25, so the project sets maven.compiler.release to 25. The slf4j-nop dependency silences the SLF4J warning that Smile prints when no logging backend is present.

6.1. Smile

Smile takes the features as a double[][] array and the labels as an int[] array. The static fit() method trains the model, and predict() returns the class index.

import smile.classification.LogisticRegression;

public class SmileExample {

  public static void main(String[] args) {
    // Features: weight in grams, length in cm. Labels: 0 = apple, 1 = banana
    double[][] x = {
        {150, 7}, {170, 8}, {140, 7}, {160, 8},
        {120, 18}, {130, 20}, {110, 17}, {125, 19}};
    int[] y = {0, 0, 0, 0, 1, 1, 1, 1};

    LogisticRegression model = LogisticRegression.fit(x, y);

    int fruit1 = model.predict(new double[] {155, 7});
    int fruit2 = model.predict(new double[] {118, 19});
    System.out.println("155 g, 7 cm  -> " + (fruit1 == 0 ? "apple" : "banana"));
    System.out.println("118 g, 19 cm -> " + (fruit2 == 0 ? "apple" : "banana"));
  }
}
155 g, 7 cm  -> apple
118 g, 19 cm -> banana

6.2. Tribuo

Tribuo needs more code, because every row is an Example with named features and a typed Label, and every dataset records where its data came from. In return, the model knows the names of its outputs, and its provenance tells us which trainer built it. An Example always carries a label, so the row we predict gets the placeholder label unknown, which the model ignores.

import org.tribuo.Example;
import org.tribuo.Model;
import org.tribuo.MutableDataset;
import org.tribuo.Prediction;
import org.tribuo.classification.Label;
import org.tribuo.classification.LabelFactory;
import org.tribuo.classification.sgd.linear.LogisticRegressionTrainer;
import org.tribuo.impl.ArrayExample;
import org.tribuo.provenance.SimpleDataSourceProvenance;

public class TribuoExample {

  static Example<Label> fruit(String label, double weight, double length) {
    return new ArrayExample<>(new Label(label), new String[] {"weight", "length"}, new double[] {weight, length});
  }

  public static void main(String[] args) {
    LabelFactory factory = new LabelFactory();
    MutableDataset<Label> data = new MutableDataset<>(new SimpleDataSourceProvenance("fruits", factory), factory);
    data.add(fruit("apple", 150, 7));
    data.add(fruit("apple", 170, 8));
    data.add(fruit("apple", 140, 7));
    data.add(fruit("apple", 160, 8));
    data.add(fruit("banana", 120, 18));
    data.add(fruit("banana", 130, 20));
    data.add(fruit("banana", 110, 17));
    data.add(fruit("banana", 125, 19));

    Model<Label> model = new LogisticRegressionTrainer().train(data);

    Prediction<Label> p = model.predict(fruit("unknown", 118, 19));
    System.out.println("118 g, 19 cm -> " + p.getOutput().getLabel());
    System.out.println("Trained by: " + model.getProvenance().getTrainerProvenance().getClassName());
  }
}
118 g, 19 cm -> banana
Trained by: org.tribuo.classification.sgd.linear.LogisticRegressionTrainer

Both models classify the unseen 118 g, 19 cm fruit as a banana. Smile needs about half the code, whereas Tribuo returns a label name instead of a number and keeps a record of how the model was trained.

7. Which Library to Pick for Each Situation

Most teams can decide with two questions. Where is the model trained, and what license can the product accept? The map and the table give the answer for the common cases.

Decision map for Java machine learning libraries. If the model is trained in Python, use ONNX Runtime for ONNX files, DJL for PyTorch or Hugging Face models and TensorFlow Java for SavedModels. If the model is trained in Java on data that fits in memory, use Smile when GPL or a commercial license is fine and Tribuo when an Apache license is required. If the data needs a cluster, use Spark MLlib. For learning, use Weka.
Where the model is trained decides the group of libraries, and the license or the data size decides the library.
SituationLibrary
Serve a PyTorch or Hugging Face model from a Spring Boot serviceDJL
Serve a model the data science team exported to ONNXONNX Runtime
Serve an existing TensorFlow SavedModelTensorFlow Java
Train a classifier on tabular data in an open-source projectSmile
Train a classifier in closed-source software without buying a licenseTribuo
Audit how every production model was trainedTribuo
Train on billions of rows in a data lakeSpark MLlib
Learn machine learning or compare algorithms in a GUIWeka
Add chat, RAG or tool calling with an LLMSpring AI or LangChain4j, not an ML library

The last row matters because many searches for Java machine learning are about LLMs. Calling a hosted model needs an LLM framework such as Spring AI or LangChain4j, and none of the libraries above is the right tool for that job.

8. Java Machine Learning FAQs

These answers rely on the Maven Central releases and license files from October 2026, and both change over time.

8.1. Is Java good for machine learning?

Java is a good choice for running models in production and for classic machine learning on tabular data. Python has far more libraries for research and deep learning training, so most teams train in Python and serve the model from Java with DJL or ONNX Runtime.

8.2. Can we use Smile in a commercial product?

Yes, if we release the product’s source under GPLv3 or buy a commercial license from HALO.AI. The Smile license file asks closed-source products, including internal apps and SaaS products, for a commercial license, so we ask the legal team before adding the dependency.

8.3. Is Deeplearning4j still maintained?

The last release on Maven Central is 1.0.0-M2.1 from August 2022. The repository still exists, but a library without a release for four years is a risk for a new project, so we use DJL or ONNX Runtime for deep learning inference instead.

8.4. Can we run a scikit-learn or XGBoost model in Java?

Yes. We export the model to ONNX with sklearn-onnx or onnxmltools and run it with ONNX Runtime. Tribuo can also load XGBoost models through its XGBoost module.

9. Conclusion

Choosing a Java machine learning library starts with the job. For models trained in Python, DJL and ONNX Runtime are the active, permissively licensed choices, and TensorFlow Java covers teams with existing SavedModels. For training in Java, Smile has a large set of algorithms but a GPL or commercial license, and Tribuo has an Apache license and provenance but a slower release pace. Weka remains the best place to learn, and Spark MLlib is the answer when the data needs a cluster.

Whatever we pick, we check the license and the date of the last release before the dependency reaches production, because those two facts change more often than the feature lists.

10. References

Happy Learning !!

Source Code on Github

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.