Building LLM Inference Engines in Java with TornadoVM and the NVIDIA Ecosystem
The session shows how Java can run GPU-accelerated local LLMs natively with TornadoVM, using NVIDIA libraries without JNI. It covers GPULlama3.java, quantized inference, Quarkus and LangChain4J integr...