Introduction
The rise of large language models (LLMs) has led to their widespread adoption in various industries and use cases across healthcare, finance, legal, tech, and customer support. The LLM development space has become so competitive that different organizations have begun to develop custom LLMs for various tasks, supporting multiple use cases. As a result, businesses have a pool of models to select from for their specific needs. However, the first question that arises is, how do they choose the best model?
Many factors should be considered when selecting LLMs for any use case, including model size and architecture, latency and speed, accuracy and performance, and cost and scalability, among others. For instance, in the case of model size, it is evident that larger models tend to perform better as they can capture more complex patterns and nuances in languages and have a larger number of training parameters and training data. Another example is that some companies prioritize latency, where they are less concerned with whether responses are extremely high-quality but need them quickly. These are some ways businesses compare and choose the best model for them.
However, once a model is selected, it is not guaranteed to perform optimally with your data. These models are trained on generic data, and your data could have a slightly different distribution. After choosing an LLM, there are additional ways to tailor it to your information and use case. One such approach is focusing on “hyperparameters.” You can customize hyperparameters by tuning them to ensure that the LLM meets your expectations.
In this article, you will learn about hyperparameters in LLMs and how you can optimize them to get the most out of your models.
What Are Hyperparameters in LLMs?
Like any other deep learning model, LLMs have two distinct types of parameters: model parameters and hyperparameters. Knowing how these parameters differ is essential to using them correctly.
Model parameters, such as weights and biases, are the values that are adjusted during training as the model continues to learn. The training algorithm automatically adjusts these parameters to minimize loss and improve the model’s performance on the given task.
On the other hand, hyperparameters are not learned during training; they must be manually set before the model training begins. Moreover, once the hyperparameters are set and the model is trained, you can not determine which hyperparameter values were used if you do not explicitly log them. These hyperparameters can include information like the model’s architecture, training configuration, and optimization strategy.
Hyperparameters are crucial for any machine learning or deep learning model, including LLM models, as they remove the overhead of training any model from scratch. You can use a widely recognized approach known as hyperparameter tuning to adjust the hyperparameters of an LLM model to achieve the desired performance.
Common hyperparameters in LLMs
Now, let’s check the curated list of hyperparameters that can affect an LLM’s performance.

Figure: Large Language Model Hyperparameters
Source
Model Size
The first parameter that takes precedence is the model size. Model size represents the number of layers in an LLM and, therefore, the number of parameters in a model. Generally speaking, models with a large number of parameters can handle more complex tasks. The model has many layers and weights that can learn linguistic, logical, and intricate relationships among different words and tokens. As you might be anticipating, there is also a “but”. Since these large models require a significantly larger dataset and more computational resources, they take much longer to train and are more costly. Moreover, these models are highly prone to overfitting due to their large size, and sometimes, due to small datasets, they may perform well on the training dataset but fail to generalize on the validation dataset.
A comparative timeline of major LLMs, evaluated by their MMLU (Massive Multitask Language Understanding) benchmark scores, is shown below. The vertical axis shows MMLU performance (with 70+ considered ideal and 89.8 equating to human expert level), while the horizontal axis represents model release dates from pre-2022 through mid-2025.
Smaller models (Small Language Models [SLMs]), on the other hand, have fewer layers, parameters, and weights and are a better choice for simpler tasks that do not require the identification of complex relationships in the data. Moreover, these models do not require a massive dataset for training, need less computation power, and are far easier to deploy than the huge models. One of the most significant advantages of using small models is that they can be easily customized for various use cases and further optimized through techniques such as quantization and fine-tuning.
Learning Rate
Learning rate determines how quickly a model updates its weights during training and ranges between 0 and 1 in the model training and fine-tuning process. This update relies on calculating the loss function, i.e., how often the model predicts the wrong label during the training process. A learning rate that is too high can accelerate the training process, but can cause the model to overshoot the optimal value. Conversely, a low learning rate can increase the stability and improve generalization, but may result in slow convergence and sometimes get stuck.
The figure below illustrates a typical learning rate schedule that combines a linear warmup and a half-cycle cosine decay. The learning rate starts at zero, increases linearly to a peak during the warmup phase, and then gradually decreases following a cosine function to stabilize training and improve convergence.
Ideally, the learning rate should be adjusted using a learning rate scheduler while the model training progresses. Some of the most common learning rate schedules include time-based decay, step decay, and exponential decay.
Batch Size
Batch size refers to the number of data samples that the model can process simultaneously. In the case of LLMs, a larger batch size can stabilize and speed up the training process, but requires more GPU memory. On the other hand, smaller batch sizes require less memory and compute power, and can improve the model’s ability to learn effectively from each sample in the dataset. To make it very clear, batch size is always dependent on your hardware capabilities.
The diagram below illustrates gradient accumulation with data parallel training across four GPUs. Each GPU processes two micro-batches per step and accumulates gradients locally for two steps before synchronizing via an Allreduce operation. This results in an effective batch size of 16 (2 micro-batches × 2 accumulation steps × 4 GPUs), enabling larger batch training without increasing per-GPU memory load.
Number of Epochs
An epoch is a complete pass through the dataset in the model training process. When you define the number of epochs, you instruct your model to examine the entire dataset for the specified times. If you set the number of epochs too high, you will improve the model’s learning ability, but this can lead to overfitting if not carefully monitored. If you use fewer epochs, your model may struggle to learn from the data, leading to underfitting. The goal is to find the optimal balance so your model can generalize well.
The diagram below explains the training and validation loss curves across epochs. It exhibits initial convergence, followed by overfitting (an increase in validation loss despite a decrease in training loss) and eventual recovery. This behavior suggests the influence of learning rate scheduling or regularization techniques, which improve generalization in later stages.
Attention Heads
In a transformer-based model, Attention Heads enable the model to focus on different parts of the input simultaneously. This hyperparameter allows the model to capture unique relationships, such as grammatical structure, word dependencies, or contextual meanings. Using multiple attention heads can enhance the model’s performance as it can attend to more diverse patterns. Still, it also adds to computational complexity and memory usage.
In the figure below, you can see the illustration of Scaled Dot-Product Attention (left), which computes attention using dot products between queries and keys, and Multi-Head Attention (right), which runs multiple parallel attention mechanisms to capture diverse contextual information and enhance model expressiveness in transformer architectures.
Max Output Tokens
Max Output Tokens define the maximum number of allowed tokens an LLM can generate in a single response. This hyperparameter is crucial as it works as a hard limit to control both the verbosity of the response and the resources, such as memory and compute time. Setting an appropriate value for the output tokens can help prevent the model from producing excessively long and illogical answers, which is essential in use cases like summarization, chatbots, and API-based responses.
The diagram below shows that the sum of input and max output tokens should always be less than the LLM context length.
On the other hand, if the max output tokens are set too low, it can cut off the meaningful content prematurely, leading to an incomplete and unhelpful response. That is why this hyperparameter is purely dependent on the task at hand. For example, a lower value is ideal for certain functions, such as intent detection or classification. Still, for tasks like story generation and summarization, you need a larger value for this hyperparameter.
Decoding Type
Decoding generates text from an LLM’s internal representation (vector embeddings). The decoding types control how an LLM generates the text during inference by determining the strategy for the next token at each step.

Figure: LLM Decoding
Source
Three of the most common decoding types are as follows:
- Greedy Decoding: This decoding selects the highest probable token at each step. While it is simple and fast enough, it can result in repetitive and less creative responses.
- Beam Search: This is an advanced version of greedy decoding, where it keeps track of multiple sequences (called beams) at once and explores various high-probability paths before selecting the best one. While it creates a more coherent response than greedy decoding, it is computationally more expensive.
- Sampling-Based Method: This method creates randomness by selecting tokens from a probability distribution instead of choosing the token with the highest probability. This method allows models to generate diverse and creative responses, which are helpful for tasks like storytelling or dialogue generation. The only drawback of this method is that it can produce inconsistent or less factual text if not tuned carefully.
When working with deterministic tasks, methods like greedy or beam search are preferred. For use cases where creativity or variation is important, sampling-based methods are the best option.
Top-p and Top-k Sampling
These hyperparameters are a subpart of the sampling-based decoding method. Both Top-p and Top-k control how tokens are selected, and they balance creativity and coherence.
Top-k sampling narrows the model’s choices to the top k highest-probability tokens at each step and then randomly selects the next token from this subset. This way, it avoids the low-probability words while allowing for some variation. For example, if k is set to 15, the model selects the next best token from this subset of 15 rather than looking in the entire corpus.
Top-p sampling, also known as nucleus sampling, dynamically selects a group of tokens whose cumulative probability adds up to a threshold p (eg, 0.9).
The diagram below compares Top-k and Top-p sampling methods in text generation-Top-k limits token choices to the highest k probabilities (e.g., 2). Top-p includes the smallest set of tokens whose cumulative likelihood exceeds a threshold (e.g., 0.96), enabling more adaptive and fluent text generation.
While top-k can provide precise control over the number of tokens, top-p offers more flexibility based on the probability distribution; ultimately, both aim to produce more diverse, human-like text.
Temperature
Temperature is the key parameter in the decoding methods that control the randomness or creativity of a language model’s output. The main factor influencing the range of possible output tokens is that the model is restricted from using the next most probable token from the probability distribution determined by the temperature. When the model predicts the next token, it usually assigns probabilities to each token in its vocabulary. Setting the temperature low (e.g., 0.2 or 0.5) sharpens the distribution. It ensures the model selects the next most probable token, leading to more deterministic and focused responses.
On the other hand, if you set the temperature high (e.g., 0.9 or 1.0), the model has a wider range for selecting a token, resulting in more varied and creative outputs.
Stop Sequence
Stop sequence is a hyperparameter that defines a specific word, token, or string that tells the LLM when to stop generating the response. It is another way to control the length of the LLM response, along with the other hyperparameter, max output tokens. This hyperparameter is highly useful for use cases such as code generation, dialogue systems, or API outputs, where you want the model to stop generating responses once a logical conclusion is reached. For example, in the coding task, the ‘####’ sequence can be used to end the code block generation.
Here is an overview of key LLM decoding and configuration options, including temperature, max tokens, stop sequences, penalties, streaming, and structured output modes. These options allow developers to precisely control model behavior and output formatting.
Moreover, the stop sequence provides flexibility in defining an integer value that represents the length of the output. In this numerical approach, one simply refers to stopping the sequence at the end of a sentence, and two refers to stopping at the end of a paragraph.
Frequency and Presence Penalties
Repetition in response is one of the biggest challenges faced by language models. Frequency and Presence Penalties hyperparameters reduce repetition and introduce diversity in the LLM response. The values of both of these hyperparameters vary between -2.0 and 2.0. The basic working principle is that these hyperparameters lower the probabilities of recently added tokens to the response, making them less likely to be used as upcoming tokens.
The figure below shows the effect of frequency penalty on token repetition in language models. Higher penalty values reduce the likelihood of repeating previously used words, promoting more diverse and less repetitive text generation.
What is LLM Hyperparameter Tuning?
Hyperparameter tuning in the context of LLMs refers to strategically adjusting the values of various hyperparameters before the training process to improve the model’s performance on a specific task or dataset. Unlike LLM parameters, these hyperparameters are set before or during the training process to achieve the best possible performance for your model. The only catch is that this tuning involves multiple trials and errors, as you need to try various combinations of hyperparameters, which ultimately leads to high time complexity and the excessive consumption of computational resources, especially in the case of LLMs.
Why is Hyperparameter Tuning Complex in LLMs?
While hyperparameter tuning is a popular concept in machine learning and deep learning, it is particularly challenging for LLMs due to several factors:
- Scale and Resource Demands: Training an LLM involves billions of parameters, so tuning experiments can take hours or sometimes even days of computation and expensive hardware (GPDs/TPUs). This makes the entire tuning process costly and time-consuming.
- Non-linear interactions: Hyperparameters often interact unpredictably with each other; for example, a reasonable learning rate may depend on the batch size, optimizer type, and other factors, so tuning them independently may not be a good idea.
- Evaluation Complexity: This is one of the significant issues with LLMs, as measuring their performance is not a straightforward task. While some use cases focus on accuracy and loss, others focus on diversity, factuality, and user management.
- High Dimensionality: As you have seen in the previous section, there are plenty of hyperparameters, each with its own range of values. This makes the hyperparameter search space vast, and finding the best set of hyperparameters becomes a bit complex.
Despite all these issues, tuning LLMs is relatively trivial to get the best performance for your LLM.
Popular Hyperparameter Optimization Techniques
Searching for the best set of hyperparameters requires a systematic approach and an understanding of what influence a hyperparameter can have on the LLM. In the case of LLMs, it is also recommended to research how the LLM you want to fine-tune was initially trained. Then, the easiest choice is to create a dictionary of all these hyperparameters and their possible values and try out each parameter combination on your own. While it may seem simple, this method is often impractical, as you may not have the time or resources. This is why you need to consider other specialized methods for hyperparameter tuning.
Grid Search
Grid Search is closely related to how you perform the manual search, as it extensively explores all the hyperparameter combinations within a specified range; the only catch is that it does this exploration automatically. While this method guarantees the full coverage of values, it is highly inefficient in the case of LLMs. Many hyperparameters lead to exponential growth in the number of combinations to try out as part of the tuning process. The only reason to mention this method in this article is that grid search is only efficient in cases where you have a smaller search space for LLM tuning.

Figure: Concept of Grid Search
Source
Random Search
Random search is an efficient alternative to grid search as it randomly selects a combination to try out instead of trying out all possible combinations. While this method seems slightly off the hook, it outperforms grid search in many cases where only a few parameters significantly influence the model performance. The only negative aspect of this tuning method is that it does not try all possible combinations; therefore, it does not guarantee that you will get the optimal parameter combinations that will lead to high LLM performance.

Source: Grid Search vs Random Search
Source
Bayesian Hyperparameter Optimization
Bayesian optimization employs a probabilistic model that guides the entire search process, making it particularly effective for expensive-to-evaluate models, such as LLMs. In a straightforward approach, this method models an object function using a surrogate model (usually a Gaussian process) that estimates the performance of different hyperparameter settings. Based on these estimates, Bayesian optimization selects the next set of hyperparameters to evaluate by balancing exploitation and exploration. This method is more efficient than the grid and random search methods as it guarantees better results while saving both time and resources required for the tuning process.

Figure: Bayesian Optimization
Source
Population-Based Training (PBT)
Population-based training is a dynamic hyperparameter tuning method that combines traditional model training with ideas from evolutionary algorithms. Instead of training a single model, PBT aims to maintain a population of models, each with its own set of hyperparameters. These models are usually trained in parallel and periodically, where the better-performing ones replace the underperforming models by inheriting their weights and hyperparameters with slight modifications and adjustments. This process is highly effective, allowing model parameters and hyperparameters to evolve together during training. While this method appears resource-intensive because multiple models are trained simultaneously, its dynamic refinement of hyperparameters can significantly save time.
The diagram below represents Population-Based Hyperparameter Optimization: Low-performing models exploit and explore hyperparameter configurations from higher-performing peers to iteratively improve model training performance.
Adaptive Low-Rank Adaptation (LoRA)
LoRA is one of the most efficient methods of fine-tuning language models without modifying their original parameters. Instead of working on full-weight metrics, LoRA inserts low-rank metrics into the model architecture, dramatically reducing the number of trainable parameters. In adaptive LoRA, the rank of the metrics is adjusted based on the task complexity or training feedback. The best part about this approach is that it saves computational resources and helps base LLMs retain their generalization ability.
Here is an illustration of Low-Rank Adaptation (LoRA): During training, the pretrained weight matrix W is frozen while low-rank matrices A and B (with B initialized to zero and A initialized randomly) are trained to adapt the model. The output is computed as h=Wx+BAx. After training, the adaptation is merged into a new weight matrix Wmerged=W+BA, allowing efficient fine-tuning with minimal additional parameters.
Best Practices for Hyperparameter Optimization in LLMs
Since optimizing hyperparameters for an LLM is a computationally expensive and complex task, you can’t just take a random model for your use case and start hyperparameter tuning/optimization. You need to know a few best practices to help you balance the performance gain and resource constraints while performing hyperparameter optimization.
Start with a smaller model or subset of data
Before going ahead with full-scope hyperparameter optimization using an extensive domain-specific model, you should begin with a smaller version of the model and possibly a reduced dataset. This method is helpful as it allows rapid experimentation and quick feedback, and helps identify promising hyperparameters with low computational costs. Once you have identified the optimal hyperparameters for the smaller model, those can work as a starting point for tuning the larger models. This method is quite effective as it accelerates development and avoids wasting time trying suboptimal configurations early in the development stage.
Use domain knowledge to reduce the search space
Researchers and developers have focused more on developing industry and using case-specific LLMs during the past few years. For example, an LLM trained on medical data might perform much better than a generic LLM on biomedical use cases. This is why prior knowledge about the task and model architecture can be required to narrow the hyperparameter search space and limit unnecessary resource usage. For instance, if you know that specific learning rates work well with transformer-based architectures or if a particular batch size aligns well with your hardware setup, things get pretty straightforward. You can eliminate unnecessary combinations and improve the likelihood of finding high-performing combinations more quickly.
Monitor resource usage and avoid overfitting
Since LLM training and tuning are resource-intensive processes, monitoring GPU usage, memory consumption, and training time is recommended to avoid bottlenecks or inefficiencies. Moreover, keeping an eye on your model is necessary to prevent it from overfitting. Regularly monitoring and evaluating your model’s performance on a validation set, and using techniques like early stopping or checkpointing mechanisms, helps maintain generalization and avoid unnecessary tuning.
The diagram below illustrates the end-to-end process for building, validating, and deploying LLM applications, including data preparation, model tuning, agent integration, validation with human and automated feedback, and continuous monitoring in production.
Track experiments consistently
In the case of LLMs, running multiple experiments in parallel or sequentially is common, so it is possible to lose track of the configurations you have tried and why. Using standard experimentation tracking tools like Weights & Biases, MLFlow, and Comet can help you to systematically log hyperparameter configurations, performance metrics, and training artifacts. This tracking enables reproducibility, allowing you to compare different configurations and make informed decisions throughout the hyperparameter optimization process.
Below is an example dashboard that tracks key training metrics—learning rate, training loss, validation loss, and BLEU score—for two LLM experiments.
While these best practices will help you stay ahead in the hyperparameter optimization space, you can also leverage a set of tools that can streamline hyperparameter optimization for large language models (LLMs). These famous tools include Optuna, Hyperopt, Google Vizier, W&B Sweeps, and Scikit-Optimize, among others. While Optuna, Hyperopt, and Scikit-Optimize are lightweight and flexible libraries for local optimization tasks, Google Vizier and W&B Sweeps offer advanced features like distributed tuning, experiment tracking, and cloud integration.
Conclusion
After reading this article, you now know what LLM hyperparameter optimization is and how it works. You have been introduced to all the hyperparameters an LLM model uses and how tuning them can maximize your selected LLM. Then, you have seen various hyperparameter optimization techniques that can be used to tune LLM hyperparameters, along with best practices that can be followed throughout the tuning process.
You must understand that the scope of LLM hyperparameter tuning is not limited to the methods mentioned in this article. In the upcoming days, methods such as AutoML, Neural Architecture Search (NAS), and hardware-aware tuning will also be taking place, so that you can best tune your model for your specific use case.













Yaron Friedman
Amos Rimon