Skip to main content

Cloud server with MLflow

The MLflow Virtual Machine is a preconfigured cloud server with a tool for tracking machine learning (ML) experiments. The tool allows you to compare ML models, optimize AI applications, and manage access to models and data.

The image used to deploy the server includes:

  • MLflow — a tool for logging and tracking ML experiments;

  • Docker — a platform for running containerized applications;

  • Docker Compose — a tool for running multi-container applications in Docker;

  • nginx — a web server and reverse proxy;

  • PostgreSQL — an object-relational database management system;

  • RustFS — a distributed object storage system;

  • drivers required for working with graphics processing units (GPU).

Before creating a server, read the software license agreements that are included in the image.

Functional capabilities

  • experiment logging — recording metrics, hyperparameters, and artifacts throughout the model training process;
  • automated model evaluation using tools integrated into the experiment tracking system;
  • collaborative management of the full model lifecycle via a model registry;
  • deploying models in Docker, Kubernetes, Azure ML, AWS SageMaker, and other environments.

Minimum resource requirements

vCPU count2
RAM4 GB
Boot volume40 GB
GPUNot required

Get started with MLflow

For MLflow to work, the cloud server must be accessible from the Internet. To do this, when creating a server, you need to create a private subnet and attach a public floating IP address.

  1. Create a public floating IP address.

  2. Create a cloud server with MLflow.

  3. Run MLflow.

1. Create a public floating IP address

Create a public floating IP address to make the cloud server with MLflow accessible from the internet.

Use the Create a public floating IP address section of the Public floating IP addresses guide.

2. Create a cloud server with MLflow

  1. In the Control panel, on the top menu, click Products and select Cloud Servers.

  2. Click Create server.

  3. Fill in the blocks:

  4. Check the cloud server price.

  5. Click Create.

Name and placement

  1. Enter a server name. It will be set as the hostname in the operating system.

  2. Select a location where the server will be created. The available server configurations and resource costs depend on the location. You cannot change the location after the server has been created.

Source

  1. Open the Applications tab.

  2. Select MLflow VM.

  3. Optional: if you need a different current or archived application version, select the required version in the Version field.

Configuration

Select a configuration from 2 vCPU, RAM starting from 4 GB and a boot disk size starting from 40 GB. For all lines, except Shared and Dedicated, two types of server configurations are available:

  • fixed configurations — configurations of lines with different technical specifications where the resource ratio is fixed;
  • custom configurations — configurations in which you can specify any resource ratio.

Configurations use different processors depending on the line and pool segment. You can customize the selected configuration. After the server is created, you will be able to change the configuration.

  1. Open the tab with a line.

  2. Click Fixed.

  3. Optional: you can customize the configuration if you are creating a server in a multi-AZ pool ru-6 segment or ru-3b, ru-7a, and ru-7b pool segments:

    3.1. Expand the block with the configuration settings description.

    3.2. Optional: select a processor manufacturer. Choosing a manufacturer is not available in all pools.

    3.3. Optional: if you do not want physical processor cores to be pinned to the vCPUs of the cloud server, uncheck the Dedicated cores checkbox. For more information, see the Dedicated cores instruction.

    3.4. Optional: if you want to disable Hyper-Threading for a server with dedicated cores, uncheck the Hyper-Threading (SMT) checkbox.

    3.5. Optional: if you are creating a cloud server with dedicated cores and want to place a multiprocessor server on a single NUMA node, check the Mandatory placement on a single NUMA node checkbox. You can place a server with 4 vCPUs or more on one NUMA node. If the cloud server resources cannot be placed on one node, it will not be created. For more information, see the Placement on a single NUMA node section of the Dedicated Cores instruction.

  4. Select a configuration.

  5. If both local and network volumes are available in the selected configuration, select the volume to be used as the boot disk:

    • local disk — check the Local SSD NVMe disk checkbox. A server with a local disk can only be created from images and applications;
    • network volume — do not check the Local SSD NVMe disk checkbox.

    The amount of RAM allocated to the server may be less than specified in the configuration — the operating system kernel reserves part of the memory depending on the kernel version and distribution. You can check the allocated volume on the server using the sudo dmesg | grep Memory command.

Volumes

  1. If you did not check the Local SSD NVMe disk checkbox during configuration, the first specified network volume will be used as the server's boot disk. To configure it:

    1.1. Select the network boot disk type.

    1.2. Specify the network boot disk size in GB or TB. Mind the network volume limits on the maximum size.

    1.3. If you selected the Universal v2 or SSD Fast v2 disk type, specify the total number of read and write operations in IOPS. After the disk is created, you can change the IOPS amount — reduce or increase it. The number of IOPS changes is unlimited.

  2. Optional: add an additional network volume server :

    2.1. Click Add.

    2.2. Select the network volume type.

    2.3. Specify the size of the network disk in GB or TB. Keep in mind the network volume limits on the maximum size.

    2.4. If you selected the Universal v2 or SSD Fast v2 volume type, specify the total IOPS. You can change the number of IOPS after creating the volume — increase or decrease it. The number of IOPS changes is unlimited.

    After the server is created, you will be able to attach new additional volumes.

Internet

Set up public access to the server.

The cloud server will be added to a private subnet that is connected to a cloud router with 1:1 NAT mapping and internet access. Access to and from the internet will be provided through the cloud router. The server will be accessible from the internet via a public floating IP address.

  1. In the Internet connection field, select the Public floating IP address access type.

  2. Select the public floating IP address you created in step 1.

Private network

  1. In the Subnet field, select a private subnet.

  2. Optional: in the IP address field, change the default IP address.

  3. In the Router field, select an existing router or create a new one.

    If the router is not connected to the internet, it will be automatically connected to the internet after the server is created.

Security

Select security groups to filter traffic on server ports. Traffic will be blocked without security groups. If the block is missing, traffic filtering (port security) is disabled in the server network. With traffic filtering disabled, all traffic will be allowed.

Access

  1. Add an SSH key for the project to the server for secure connection:

    1.1. If an SSH key for the project has not been added to the cloud platform, click Add SSH key, enter the key name, paste the OpenSSH public key, and click Add.

    1.2. If an SSH key for the project has been added to the cloud platform, select an existing key in the SSH key field. The SSH key is only available in the pool where it is located.

  2. Optional: in the Root password field:

    2.1. Copy the root user password — the user with unrestricted privileges for all system actions.

    2.2. Save the password in a secure place and do not share it in plain text.

Additional settings

  1. If you plan to create multiple servers and want to improve infrastructure fault tolerance, add the server to a placement group:

    1.1. To create a new group, in the Placement group field, click Create.

    1.2. Select New group and enter the group name.

    1.3. Select a placement policy on different hosts:

    • preferred — soft-anti-affinity. The system will try to place servers on different hosts. If no suitable host is available when creating the server, it will be placed on the host where it fits;
    • required — anti-affinity. Servers in the group must be located on different hosts. If no suitable host is available when creating the server, the server will not be created.

    1.4. If the group has been created, select it in the Placement group field.

  2. To add additional information or filter servers in the list, add server tags. Operating system and configuration tags are added automatically. To add a new tag, enter it in the Tags field.

  3. To add a script that will be executed using the cloud-init agent when the operating system starts for the first time, in the Automation block, in the User data:

    • open the Text tab and paste the script as text;
    • or open the File tab and upload the file with the script.

3. Run MLflow

  1. Open the page http://<ip_address> in your browser.

    Specify <ip_address> — the public floating IP address of the cloud server. You can copy it in the control panel: from the top menu, click ProductsCloud Servers → in the server card, click next to the public floating IP address.

  2. Enter the username — admin.

  3. Enter the password — the server UUID. You can copy it in the control panel: from the top menu, click ProductsCloud Servers → in the server card, click next to the UUID.

  4. Click Log in.