For classic machine learning on tabular data in Java, the strongest libraries in 2026 are Smile and Tribuo, and for running deep learning models that were trained in Python, the standard choices are DJL and ONNX Runtime. Weka still fits teaching and quick experiments, and Spark MLlib fits data that does not fit on one machine.
A Java team picks a machine learning library for one of two jobs. Either we train a model inside the JVM on our own data, or a data science team trains the model in Python and we serve it from a Java service. The two jobs need different libraries, and many top-10 lists mix them together without saying so.
Training a model in Java can take a few lines. The next snippet trains a logistic regression classifier with Smile on an array of features and labels.
// Features: weight in grams, length in cm. Labels: 0 = apple, 1 = banana
double[][] x = {
{150, 7}, {170, 8}, {140, 7}, {160, 8},
{120, 18}, {130, 20}, {110, 17}, {125, 19}};
int[] y = {0, 0, 0, 0, 1, 1, 1, 1};
LogisticRegression model = LogisticRegression.fit(x, y);
int fruit1 = model.predict(new double[] {155, 7}); // 0, an apple
In the rest of the article, we compare ten Java machine learning libraries by the latest version, the license, the minimum Java version and whether the project still ships releases. After that, we train the same classifier with Smile and Tribuo, and we finish with a table that maps common situations to a library. Readers who already know their situation can jump to that table in section 7.
1. Java Machine Learning Libraries Compared
We checked every version and release date against Maven Central in October 2026. A library counts as active when it had a release in the last 12 months, because a stalled project stops receiving fixes for new Java versions and native platforms.
| Library | Job | Best for | Latest version (date) | License | Min Java | Status |
|---|---|---|---|---|---|---|
| Smile | Train | Classic ML, statistics and NLP | 6.3.0 (Aug 2026) | GPLv3 or commercial | 25 | Active |
| Tribuo | Train | Classic ML with tracked model provenance | 4.3.2 (Apr 2025) | Apache-2.0 | 8 (17 for some modules) | No release for 18 months |
| Weka | Train | Teaching, GUI experiments, classic ML | 3.8.7 (Jul 2026) | GPL-3.0 | Not stated | Active |
| DJL | Inference | Deep learning inference with PyTorch, TensorFlow or ONNX engines | 0.38.0 (Sep 2026) | Apache-2.0 | 11 | Active |
| ONNX Runtime | Inference | Fast inference for ONNX models | 1.31.0 (Oct 2026) | MIT | 8 | Active |
| TensorFlow Java | Inference | Running TensorFlow SavedModels | 1.2.0 (Jul 2026) | Apache-2.0 | 11 | Active |
| Spark MLlib | Train on a cluster | Distributed training on a cluster | 4.2.0 (Jul 2026) | Apache-2.0 | 17, 21 or 25 | Active |
| Deeplearning4j | Avoid | Deep learning on the JVM | 1.0.0-M2.1 (Aug 2022) | Apache-2.0 | Not stated | No release for 4 years |
| Jlama | Avoid | Running LLMs in pure Java | 0.8.4 (Jan 2025) | Apache-2.0 | 21 (preview features) | No release for 21 months |
| MOA | Avoid | Stream mining and concept drift | 2024.07.0 (Jul 2024) | GPL-3.0 | 8 | No release for 27 months |
Three libraries in the table use GPL-3.0, and Smile also sells a commercial license, so we check the license before we add any of them to closed-source software. The Apache-2.0 and MIT libraries have no such condition.
2. Libraries for Training Classic Models
Classic machine learning means algorithms such as logistic regression, random forests, gradient boosting and k-means on rows of numbers or categories. These models train in seconds on a laptop and are often the right first model for fraud scores or churn prediction.
2.1. Smile
Smile (Statistical Machine Intelligence and Learning Engine) has one of the largest sets of algorithms in Java, from classification and regression to clustering, manifold learning, NLP and plotting. It is fast, and its API takes plain Java arrays, so we can start without learning a data model first.
| Fact | Details |
|---|---|
| Maven coordinates | com.github.haifengl:smile-core |
| Latest version | 6.3.0, released August 2026 |
| Minimum Java | 25 for Smile 5 and later (Smile 4 needs 21) |
| License | GPLv3 or a commercial license from HALO.AI |
| Strengths | A large set of algorithms, good speed, a Scala and Kotlin API, a DataFrame class and a shell |
| Limits | The license rules out closed-source use without a paid license, and the Java 25 baseline rules out older runtimes |
The Smile license file asks for a commercial license from any closed-source use that does not release its source under GPLv3, and it names internal enterprise apps and SaaS products as examples. The Maven POM only says GNU 3, so a dependency scanner may not flag the commercial terms.
Pick Smile if the project is open source or the company buys a license, and the team wants many algorithms behind one simple API.
2.2. Tribuo
Tribuo is a machine learning library from Oracle Labs. Every Tribuo model records its provenance, which is the dataset and the trainer with all hyperparameters used to build it, so we can audit a model in production or retrain it the same way a year later.
| Fact | Details |
|---|---|
| Maven coordinates | org.tribuo:tribuo-all, or single modules such as tribuo-classification-sgd |
| Latest version | 4.3.2, released April 2025 (the main branch is on 5.0.0-SNAPSHOT) |
| Minimum Java | 8 for the core, 17 for the model card modules |
| License | Apache-2.0 |
| Strengths | Typed outputs (Label, Regressor), provenance, wrappers for XGBoost, ONNX Runtime and TensorFlow |
| Limits | No release in 18 months and fewer algorithms than Smile |
Pick Tribuo if the model must be auditable or the legal team needs an Apache license, and the team accepts the slower release pace.
2.3. Weka
Weka from the University of Waikato is the oldest library in the list, and the Data Mining textbook by Witten, Frank, Hall and Pal is built around it. Its Explorer GUI loads a CSV or ARFF file and compares algorithms without any code, and the same algorithms are available as a Java API.
| Fact | Details |
|---|---|
| Maven coordinates | nz.ac.waikato.cms.weka:weka-stable |
| Latest version | 3.8.7, released July 2026 (the development line is 3.9.7) |
| Minimum Java | Not stated on the download page, and the installers bundle OpenJDK 25 |
| License | GPL-3.0 |
| Strengths | A GUI for experiments, plus decades of examples and books |
| Limits | GPL license, and an older API built around the Instances class |
Some articles link Weka to weka.io, which is an unrelated storage company. The library lives at the University of Waikato site linked above.
Pick Weka if we are learning machine learning or want to compare algorithms in a GUI before writing code.
3. Libraries for Running Deep Learning Models
Most deep learning models are trained in Python with PyTorch or TensorFlow. The Java side loads the trained file and runs inference, so these libraries are thin layers over a native engine written in C++. Each one downloads platform-specific native binaries, which we must account for in Docker images and on ARM servers.
3.1. Deep Java Library (DJL)
DJL is an engine-agnostic deep learning library from AWS. We write the Java code once against DJL’s API, and DJL runs the model on PyTorch, TensorFlow or ONNX Runtime underneath, which makes it the most flexible choice for serving models in Java.
| Fact | Details |
|---|---|
| Maven coordinates | ai.djl:api plus an engine such as ai.djl.pytorch:pytorch-engine |
| Latest version | 0.38.0, released September 2026 |
| Minimum Java | 11 |
| License | Apache-2.0 |
| Strengths | Several engines behind one API, a model zoo, Hugging Face tokenizers, DJL Serving for model servers |
| Limits | Still on 0.x versions, and the native engines add hundreds of megabytes |
Pick DJL if we serve PyTorch or Hugging Face models from Java and want the option to switch engines later.
3.2. ONNX Runtime for Java
ONNX Runtime from Microsoft runs models saved in the ONNX format, which PyTorch, TensorFlow, scikit-learn and XGBoost can all export. The Java API is small, with an OrtEnvironment, an OrtSession per model and tensors for inputs and outputs.
| Fact | Details |
|---|---|
| Maven coordinates | com.microsoft.onnxruntime:onnxruntime, or onnxruntime_gpu for CUDA |
| Latest version | 1.31.0, released October 2026 |
| Minimum Java | 8 |
| License | MIT |
| Strengths | Fast CPU inference, GPU support, one format for models from many Python frameworks |
| Limits | Inference only, and we write the pre- and post-processing ourselves |
Pick ONNX Runtime if the data science team can export to ONNX and we want the smallest dependency for fast inference.
3.3. TensorFlow Java
TensorFlow Java is the official Java binding for TensorFlow, maintained by the SIG JVM group. It loads a TensorFlow SavedModel and can also train, though few teams train in Java with it.
| Fact | Details |
|---|---|
| Maven coordinates | org.tensorflow:tensorflow-core-platform |
| Latest version | 1.2.0, released July 2026 |
| Minimum Java | 11 |
| License | Apache-2.0 |
| Strengths | Runs SavedModels without conversion, with the full TensorFlow op set |
| Limits | The last Windows x86_64 build is 1.1.0, macOS x86_64 builds ended before 1.1.0, and the README lags behind the releases |
Pick TensorFlow Java if the models are TensorFlow SavedModels and converting them to ONNX is not an option.
4. Spark MLlib for Data That Needs a Cluster
Spark MLlib trains models on DataFrames spread across a Spark cluster. It has the classic algorithms, along with pipelines for feature engineering, and it runs the same code on a laptop or on hundreds of nodes. The DataFrame API in the spark.ml package is the one to use, because the older RDD API is in maintenance mode.
| Fact | Details |
|---|---|
| Maven coordinates | org.apache.spark:spark-mllib_2.13 |
| Latest version | 4.2.0, released July 2026 |
| Minimum Java | 17, and Spark 4.2 also supports 21 and 25 |
| License | Apache-2.0 |
| Strengths | Scales to billions of rows, pipelines, model persistence, a Java API |
| Limits | Local mode works for development, but Spark only pays off on a cluster, and its overhead is not worth it for data that fits in memory |
Pick Spark MLlib if the training data is already in a data lake or does not fit on one machine.
5. Libraries to Avoid for New Projects
Older lists still recommend libraries that have not had a release for years. They work with the Java versions they were built for, but nobody fixes their bugs or security issues, and they fall behind new JDKs and CPU architectures.
| Library | Last release | Use instead |
|---|---|---|
| Deeplearning4j | 1.0.0-M2.1, August 2022 | DJL or ONNX Runtime for inference |
| Jlama | 0.8.4, January 2025 | A hosted LLM through Spring AI or LangChain4j, or a local model server such as Ollama |
| MOA | 2024.07.0, July 2024 | Retraining a Smile or Tribuo model on a schedule. For true online learning, MOA still works, but it gets no fixes |
Deeplearning4j’s README still says the project is in active development, but no build has reached Maven Central since 2022. If a new release appears, it is worth a second look.
6. Training the Same Classifier with Smile and Tribuo
The following example is a fruit classifier that tells apples from bananas by weight in grams and length in centimeters. We build it twice, with Smile 6.3.0 and Tribuo 4.3.2 on Java 25, because the two libraries show the trade-off between a short API and a typed one.
<dependency>
<groupId>com.github.haifengl</groupId>
<artifactId>smile-core</artifactId>
<version>6.3.0</version>
</dependency>
<dependency>
<groupId>org.tribuo</groupId>
<artifactId>tribuo-classification-sgd</artifactId>
<version>4.3.2</version>
</dependency>
<dependency>
<groupId>org.slf4j</groupId>
<artifactId>slf4j-nop</artifactId>
<version>2.0.17</version>
</dependency>
Smile 6 is compiled for Java 25, so the project sets maven.compiler.release to 25. The slf4j-nop dependency silences the SLF4J warning that Smile prints when no logging backend is present.
6.1. Smile
Smile takes the features as a double[][] array and the labels as an int[] array. The static fit() method trains the model, and predict() returns the class index.
import smile.classification.LogisticRegression;
public class SmileExample {
public static void main(String[] args) {
// Features: weight in grams, length in cm. Labels: 0 = apple, 1 = banana
double[][] x = {
{150, 7}, {170, 8}, {140, 7}, {160, 8},
{120, 18}, {130, 20}, {110, 17}, {125, 19}};
int[] y = {0, 0, 0, 0, 1, 1, 1, 1};
LogisticRegression model = LogisticRegression.fit(x, y);
int fruit1 = model.predict(new double[] {155, 7});
int fruit2 = model.predict(new double[] {118, 19});
System.out.println("155 g, 7 cm -> " + (fruit1 == 0 ? "apple" : "banana"));
System.out.println("118 g, 19 cm -> " + (fruit2 == 0 ? "apple" : "banana"));
}
}
155 g, 7 cm -> apple
118 g, 19 cm -> banana
6.2. Tribuo
Tribuo needs more code, because every row is an Example with named features and a typed Label, and every dataset records where its data came from. In return, the model knows the names of its outputs, and its provenance tells us which trainer built it. An Example always carries a label, so the row we predict gets the placeholder label unknown, which the model ignores.
import org.tribuo.Example;
import org.tribuo.Model;
import org.tribuo.MutableDataset;
import org.tribuo.Prediction;
import org.tribuo.classification.Label;
import org.tribuo.classification.LabelFactory;
import org.tribuo.classification.sgd.linear.LogisticRegressionTrainer;
import org.tribuo.impl.ArrayExample;
import org.tribuo.provenance.SimpleDataSourceProvenance;
public class TribuoExample {
static Example<Label> fruit(String label, double weight, double length) {
return new ArrayExample<>(new Label(label), new String[] {"weight", "length"}, new double[] {weight, length});
}
public static void main(String[] args) {
LabelFactory factory = new LabelFactory();
MutableDataset<Label> data = new MutableDataset<>(new SimpleDataSourceProvenance("fruits", factory), factory);
data.add(fruit("apple", 150, 7));
data.add(fruit("apple", 170, 8));
data.add(fruit("apple", 140, 7));
data.add(fruit("apple", 160, 8));
data.add(fruit("banana", 120, 18));
data.add(fruit("banana", 130, 20));
data.add(fruit("banana", 110, 17));
data.add(fruit("banana", 125, 19));
Model<Label> model = new LogisticRegressionTrainer().train(data);
Prediction<Label> p = model.predict(fruit("unknown", 118, 19));
System.out.println("118 g, 19 cm -> " + p.getOutput().getLabel());
System.out.println("Trained by: " + model.getProvenance().getTrainerProvenance().getClassName());
}
}
118 g, 19 cm -> banana
Trained by: org.tribuo.classification.sgd.linear.LogisticRegressionTrainer
Both models classify the unseen 118 g, 19 cm fruit as a banana. Smile needs about half the code, whereas Tribuo returns a label name instead of a number and keeps a record of how the model was trained.
7. Which Library to Pick for Each Situation
Most teams can decide with two questions. Where is the model trained, and what license can the product accept? The map and the table give the answer for the common cases.

| Situation | Library |
|---|---|
| Serve a PyTorch or Hugging Face model from a Spring Boot service | DJL |
| Serve a model the data science team exported to ONNX | ONNX Runtime |
| Serve an existing TensorFlow SavedModel | TensorFlow Java |
| Train a classifier on tabular data in an open-source project | Smile |
| Train a classifier in closed-source software without buying a license | Tribuo |
| Audit how every production model was trained | Tribuo |
| Train on billions of rows in a data lake | Spark MLlib |
| Learn machine learning or compare algorithms in a GUI | Weka |
| Add chat, RAG or tool calling with an LLM | Spring AI or LangChain4j, not an ML library |
The last row matters because many searches for Java machine learning are about LLMs. Calling a hosted model needs an LLM framework such as Spring AI or LangChain4j, and none of the libraries above is the right tool for that job.
8. Java Machine Learning FAQs
These answers rely on the Maven Central releases and license files from October 2026, and both change over time.
8.1. Is Java good for machine learning?
Java is a good choice for running models in production and for classic machine learning on tabular data. Python has far more libraries for research and deep learning training, so most teams train in Python and serve the model from Java with DJL or ONNX Runtime.
8.2. Can we use Smile in a commercial product?
Yes, if we release the product’s source under GPLv3 or buy a commercial license from HALO.AI. The Smile license file asks closed-source products, including internal apps and SaaS products, for a commercial license, so we ask the legal team before adding the dependency.
8.3. Is Deeplearning4j still maintained?
The last release on Maven Central is 1.0.0-M2.1 from August 2022. The repository still exists, but a library without a release for four years is a risk for a new project, so we use DJL or ONNX Runtime for deep learning inference instead.
8.4. Can we run a scikit-learn or XGBoost model in Java?
Yes. We export the model to ONNX with sklearn-onnx or onnxmltools and run it with ONNX Runtime. Tribuo can also load XGBoost models through its XGBoost module.
9. Conclusion
Choosing a Java machine learning library starts with the job. For models trained in Python, DJL and ONNX Runtime are the active, permissively licensed choices, and TensorFlow Java covers teams with existing SavedModels. For training in Java, Smile has a large set of algorithms but a GPL or commercial license, and Tribuo has an Apache license and provenance but a slower release pace. Weka remains the best place to learn, and Spark MLlib is the answer when the data needs a cluster.
Whatever we pick, we check the license and the date of the last release before the dependency reaches production, because those two facts change more often than the feature lists.
10. References
- Smile
- Tribuo Documentation
- Weka
- DJL Documentation
- ONNX Runtime Java API
- TensorFlow Java
- Spark MLlib Guide
- Maven Central
Happy Learning !!