> For the complete documentation index, see [llms.txt](https://netmind-power.gitbook.io/netmind-power-documentation/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://netmind-power.gitbook.io/netmind-power-documentation/inference/dedicated-endpoints.md).

# Dedicated Endpoints

NetMind also offers serverless inference capabilities. With this feature, you are billed based on the actual runtime of the inference service, which can significantly reduce operational costs for small to medium-sized applications or for applications with peak-and-valley API usage patterns.

1. To create a dedicated endpoint, go to " Dedicated Endpoint" page. Click "+ New" button

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2FOYVBGVjhYVTA7Oqx4fGx%2Fimage.png?alt=media&amp;token=c6484b83-62f7-46be-87b5-1897f5c8f8c0" alt=""><figcaption><p>Dedicated Endpoint page</p></figcaption></figure>

1. you need to choose either create an endpoint with "Custom Image" or "NetMind Model". If you choose "Custom Image" you need to then select an appropriate GPU, which depends on the VRAM requirements of the model and the inference efficiency you expect.

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2FJoDY87vZIdbgf7UsgfRZ%2Fimage.png?alt=media&amp;token=3c756fff-91e6-44df-8814-48783428b6e9" alt=""><figcaption><p>Dedicated Endpoint page</p></figcaption></figure>

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2Fe2lSc1s5Ej4zvYHBVYys%2Fimage.png?alt=media&amp;token=cae39345-44dd-4088-a1b4-1787007f581f" alt=""><figcaption><p>Select inference GPU</p></figcaption></figure>

3. Next, you need to fill in a series of information, including basic details such as the inference name and description, payment method, and scaling method. The platform currently supports both manual and automatic scaling.

   * **Manual Scaling**: Users need to manually manage the number of workers for model inference instances.
   * **Automatic Scaling**: The platform adjusts the number of model inference instances automatically based on the scaling rules chosen by the user.

   The platform supports two types of scaling rules: **concurrency** and **RPS (Requests Per Second)**. Both rules allow users to configure a **threshold**. When the platform detects that the request volume exceeds the threshold based on the selected rule, it increases the number of model inference workers. Conversely, if the request volume falls below the threshold, it decreases the workers.

   Additionally, in automatic mode, the platform allows setting a maximum or minimum number of workers:

   * Setting the **minimum replica count greater than 0** can avoid cold start issues.
   * Setting a **maximum replica count** helps users prevent exceeding their budget.

   Additionally, you need to package your inference service into a Docker image and push it to a container registry. After that, enter the image's url in the provided form. We support both public and private images.

{% hint style="info" %}
the maximum number of workers is also limited by the number of machines currently available on the platform.
{% endhint %}

{% hint style="info" %}
Please note that the inference service inside the image must be started on port 8080, as we currently do not support inference services running on other ports.
{% endhint %}

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2FvNLIBE2ddbud7sk9nyHU%2Fimage.png?alt=media&amp;token=f0504d95-caf2-495a-97cb-cd1e9abe6361" alt=""><figcaption></figcaption></figure>

4. Next, you can see a newly created serverless instance with status "**Deploying**". You can click **"Detail"** to view more information. Here, you can see the details you provided during creation, obtain a model inference endpoint, check billing information, monitor the number of workers and access the running logs of the workers. Additionally, you can stop or edit this Serverless instance from this page.&#x20;

   Serverless instance will have these statuses:

   * **Deploying**: The status immediately after creating a new instance or redeploying a stopped instance. It means the instance is being initializing.
   * **Running**: Typically follows the "Deploying" status. This indicates the instance is running normally, can be scaled, and the endpoint is accessible.
   * **Stopped**: The instance is stopped, and the number of workers will be reduced to zero.
   * **Failed**: The instance deployment or scaling failed due to an error, which may cause the endpoint to become inaccessible.

   <figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2F98qhu7gmZFRp5fbydAfd%2Fimage.png?alt=media&amp;token=4d0423bf-ce3f-4956-9f5c-380f76fe5407" alt=""><figcaption></figcaption></figure>

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2FtqebhmEDre0Qp3Bnqp6d%2Fimage.png?alt=media&amp;token=c7586a11-c89f-4d00-8719-7221a59146ad" alt=""><figcaption></figcaption></figure>

4. Finally, you can obtain the request URL by clicking **Endpoint Info** or checking the **Request URL** displayed on the interface.

<figure><img src="https://2171129615-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fzj4sv2NjN42aoPb00bj8%2Fuploads%2Fy61uCQUbMMoVjfBmUzll%2Fimage.png?alt=media&amp;token=beb3b78e-20d0-4852-a69c-ce9265126572" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Note that when making actual inference requests, you may need to concatenate the platform-provided URL with the URL path required by the inference service in the image, depending on the service's requirements.

When making requests to the  serverless endpoint, you still need to use an API token. This is to ensure the security of the API.
{% endhint %}
