Post

Training a CNN Model for End to End Visual Control and Fine-Tuning Visual Human Detection

For my first two internship tasks, I had to work on two exercises in Unibotics: End to End Visual Control and Visual Object Detection.

End to End Visual Control. Creating the CNN and training the model

In this exercise, a car must complete a lap by following a red line. The exercise documentation requires using an AI model trained with a CNN (Convolutional Neural Network).

Because of this, I first needed to create and train a neural network using the base datasets provided by the exercise. I chose the dataset for the Simple Circuit.

I decided to use Python with PyTorch. Since it was my first time using PyTorch, I followed the PyTorch documentation and also researched how CNNs work and how to build a CNN.

I also had to balance the dataset. Many of the images show a straight line, which means the car’s angular velocity was zero most of the time. To prevent the model from learning to predict only zero, I created a balancing system that equalizes the images used in the training process.

After some days of programming and testing, I arrived at the following code for training the neural network.

FollowLineDataset.py

This file contains the FollowLineDataset class. I created it to keep the code better organized and to take advantage of PyTorch’s data utilities.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
import torch
import torch.nn as nn
import torch.optim as optim
import os
from PIL import Image
import numpy as np

class FollowLineDataset(torch.utils.data.Dataset):

    def __init__(self, csvfile):
        super().__init__()

        self.csvfile=csvfile

        self.data=[]

        with open(self.csvfile, "r", encoding="utf-8") as f:
                lines=f.readlines()

        lines=lines[1:]

        for line in lines:
             image_name, v, w = line.strip().split(",")
             self.data.append(
                  (
                       image_name,
                       float(v),
                       float(w)
                  )
             )

        dataset_folder = os.path.dirname(csvfile)
        self.image_paths = {}

        for folder in os.listdir(dataset_folder):

            if csvfile.endswith("train.csv") and not folder.startswith("train_images_part_"):
                continue

            if csvfile.endswith("validation.csv") and not folder.startswith("test_images"):
                continue

            folder_path = os.path.join(
                dataset_folder,
                folder
            )

            if not os.path.isdir(folder_path):
                continue

            for file in os.listdir(folder_path):

                if file.endswith(".png"):

                    image_path = os.path.join(
                        folder_path,
                        file
                    )

                    idx=os.path.join(folder, file)
                    self.image_paths[idx] = image_path

    def __len__(self):
         return len(self.data)

    def __getitem__(self, index):
         image_name, v, w = self.data[index]
         image_path = self.image_paths[image_name]
         image = Image.open(image_path)
         image = np.array(image)
         image = image / 255

         image = torch.tensor(image, dtype=torch.float32)
         image = image.permute(2,0,1)

         target=torch.tensor([v, w], dtype=torch.float32)

         return image, target

FollowLineCNN.py

This file contains the FollowLineCNN class, which defines the custom CNN I used to train the model. I decided to use only three convolutional layers, changing a few parameters along the way for testing.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
import torch
import torch.nn as nn
import torch.optim as optim

class FollowLineCNN(nn.Module):

    def __init__(self):
        super().__init__()

        self.conv1=nn.Conv2d(3, 16, kernel_size=3, stride=1, padding=1)
        self.relu1=nn.ReLU()
        self.pool1=nn.MaxPool2d(kernel_size=2, stride=2)

        self.conv2=nn.Conv2d(16, 16, kernel_size=3, stride=1, padding=1)
        self.relu2=nn.ReLU()
        self.pool2=nn.MaxPool2d(kernel_size=2, stride=2)

        self.conv3=nn.Conv2d(16, 16, kernel_size=3, stride=1, padding=1)
        self.relu3=nn.ReLU()
        self.pool3=nn.MaxPool2d(kernel_size=2, stride=2)

        self.flatten=nn.Flatten()
        self.linear=nn.Linear(76800, 2)

    def forward(self, x):

        x = self.conv1(x)
        x = self.relu1(x)
        x = self.pool1(x)

        x = self.conv2(x)
        x = self.relu2(x)
        x = self.pool2(x)

        x = self.conv3(x)
        x = self.relu3(x)
        x = self.pool3(x)

        x = self.flatten(x)
        x = self.linear(x)

        return x

model_training_ai_follow_line.py

This is the main script for training and exporting the model. To balance the dataset, I used a weight-based system: each sample is weighted by the inverse frequency of its class, so images with an angular velocity of zero (the most common) get less weight and images with non-zero angular velocity get more. I exported the model to .onnx so it can be used inside the Unibotics platform, as the exercise statement indicates.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
import os
import matplotlib.pyplot as plt
import numpy as np

from PIL import Image
import torch
import torch.nn as nn
import torch.optim as optim

from FollowLineCNN import FollowLineCNN
from FollowLineDataset import FollowLineDataset


def datainfo(n_train_images, n_validation_images, v_values, w_values):
    print(f"Total number of images: {len(n_train_images) + len(n_validation_images)}")
    print(f"Total number of training samples: {len(n_train_images)}")
    print(f"Min v: {min(v_values)}")
    print(f"Max v: {max(v_values)}")
    print(f"Mean v: {sum(v_values) / len(v_values):.4f}")
    print(f"Min w: {min(w_values)}")
    print(f"Max w: {max(w_values)}")
    print(f"Mean w: {sum(w_values) / len(w_values):.4f}\n")


def hist_v_w(v_values, w_values):
    plt.figure(figsize=(10, 4))

    plt.subplot(1, 2, 1)
    plt.hist(w_values, bins=10, color="skyblue", density=True)
    plt.xlabel("Value")
    plt.ylabel("Frequency")
    plt.title("W Histogram")

    plt.subplot(1, 2, 2)
    plt.hist(v_values, bins=10, color="skyblue", density=True)
    plt.xlabel("Value")
    plt.ylabel("Frequency")
    plt.title("V Histogram")

    plt.tight_layout()
    plt.show()


def diagram_v_w(v_values, w_values):
    plt.scatter(v_values, w_values)
    plt.colorbar()
    plt.show()


def dataimg():
    dataset_path = "/media/hectormc/USB/Datasets_Jderobot/follow_line_dataset"
    for actual_folder, _, files in os.walk(dataset_path):
        if actual_folder == os.path.join(dataset_path, "test_images"):
            continue
        for file in files:
            if not file.endswith(".png"):
                continue
            rute = os.path.join(actual_folder, file)
            img = Image.open(rute)
            print(img.size, img.mode)


def create_sample_weights(dataset):
    w_values = np.array([sample[2] for sample in dataset.data], dtype=np.float32)
    classes = np.digitize(w_values, bins=[-1.0, -0.2, 0.2, 1.0])
    class_counts = np.bincount(classes, minlength=5)

    print("\nClass distribution:")
    for i, count in enumerate(class_counts):
        print(f"Class {i}: {count} samples")

    class_weights = 1.0 / class_counts
    print("\nClass weights:")
    for i, weight in enumerate(class_weights):
        print(f"Class {i}: {weight:.8f}")

    sample_weights = class_weights[classes]
    return torch.tensor(sample_weights, dtype=torch.double)


def main():
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    print(f"Device used: {device}")
    if torch.cuda.is_available():
        print(f"GPU: {torch.cuda.get_device_name(0)}")

    base_path = "/media/hectormc/USB/Datasets_Jderobot/follow_line_dataset"
    train_dataset = FollowLineDataset(os.path.join(base_path, "train.csv"))
    validation_dataset = FollowLineDataset(os.path.join(base_path, "validation.csv"))

    sample_weights = create_sample_weights(train_dataset)
    sampler = torch.utils.data.WeightedRandomSampler(
        weights=sample_weights,
        num_samples=len(train_dataset),
        replacement=True,
    )

    train_loader = torch.utils.data.DataLoader(
        train_dataset, batch_size=16, sampler=sampler, num_workers=4
    )
    validation_loader = torch.utils.data.DataLoader(
        validation_dataset, batch_size=16, shuffle=False, num_workers=4
    )

    model = FollowLineCNN().to(device)
    loss_function = nn.MSELoss()
    optimizer = optim.Adam(model.parameters(), lr=0.001)

    n_epoch = 15

    for epoch in range(n_epoch):
        model.train()
        train_loss = 0.0
        for images, targets in train_loader:
            images = images.to(device)
            targets = targets.to(device)

            predictions = model(images)
            loss = loss_function(predictions, targets)

            optimizer.zero_grad()
            loss.backward()
            optimizer.step()

            train_loss += loss.item()

        train_loss /= len(train_loader)

        model.eval()
        validation_loss = 0.0
        with torch.no_grad():
            for images, targets in validation_loader:
                images = images.to(device)
                targets = targets.to(device)

                predictions = model(images)
                loss = loss_function(predictions, targets)
                validation_loss += loss.item()

        validation_loss /= len(validation_loader)

        print(
            f"Epoch: {epoch + 1}/{n_epoch} | "
            f"Train Loss: {train_loss:.4f} | "
            f"Validation Loss: {validation_loss:.4f}"
        )

    # Export to ONNX
    model.eval()
    model_cpu = model.to("cpu")
    dummy_input = torch.randn(1, 3, 480, 640)
    torch.onnx.export(
        model_cpu,
        dummy_input,
        "follow_line.onnx",
        input_names=["image"],
        output_names=["output"],
    )
    print("\nLast epoch model exported:\nfollow_line.onnx")


if __name__ == "__main__":
    main()

The last step to complete this task is testing the trained model in the exercise in Unibotics.

Visual Object Detection. Downloading the pre-trained model, the person dataset, and fine-tuning

While working on End to End Visual Control, I was also working on the Visual Object Detection exercise. Its objective is to detect people in front of the webcam and draw a bounding box around each one. To achieve this, the exercise requires fine-tuning a pre-existing model by following the guide for fine-tuning pre-existing models in PyTorch.

Over the last week I followed that fine-tuning tutorial. I downloaded the pre-trained model mobilenet-v1-ssd and part of the person dataset (only a subset of the images, because my device did not have enough space for the whole dataset). I am now fine-tuning the MobileNet model with the downloaded subset of the person dataset.

This post is licensed under CC BY 4.0 by the author.