Training a CNN Model for End to End Visual Control and Fine-Tuning Visual Human Detection
For my first two internship tasks, I had to work on two exercises in Unibotics: End to End Visual Control and Visual Object Detection.
End to End Visual Control. Creating the CNN and training the model
In this exercise, a car must complete a lap by following a red line. The exercise documentation requires using an AI model trained with a CNN (Convolutional Neural Network).
Because of this, I first needed to create and train a neural network using the base datasets provided by the exercise. I chose the dataset for the Simple Circuit.
I decided to use Python with PyTorch. Since it was my first time using PyTorch, I followed the PyTorch documentation and also researched how CNNs work and how to build a CNN.
I also had to balance the dataset. Many of the images show a straight line, which means the car’s angular velocity was zero most of the time. To prevent the model from learning to predict only zero, I created a balancing system that equalizes the images used in the training process.
After some days of programming and testing, I arrived at the following code for training the neural network.
FollowLineDataset.py
This file contains the FollowLineDataset class. I created it to keep the code better organized and to take advantage of PyTorch’s data utilities.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
import torch
import torch.nn as nn
import torch.optim as optim
import os
from PIL import Image
import numpy as np
class FollowLineDataset(torch.utils.data.Dataset):
def __init__(self, csvfile):
super().__init__()
self.csvfile=csvfile
self.data=[]
with open(self.csvfile, "r", encoding="utf-8") as f:
lines=f.readlines()
lines=lines[1:]
for line in lines:
image_name, v, w = line.strip().split(",")
self.data.append(
(
image_name,
float(v),
float(w)
)
)
dataset_folder = os.path.dirname(csvfile)
self.image_paths = {}
for folder in os.listdir(dataset_folder):
if csvfile.endswith("train.csv") and not folder.startswith("train_images_part_"):
continue
if csvfile.endswith("validation.csv") and not folder.startswith("test_images"):
continue
folder_path = os.path.join(
dataset_folder,
folder
)
if not os.path.isdir(folder_path):
continue
for file in os.listdir(folder_path):
if file.endswith(".png"):
image_path = os.path.join(
folder_path,
file
)
idx=os.path.join(folder, file)
self.image_paths[idx] = image_path
def __len__(self):
return len(self.data)
def __getitem__(self, index):
image_name, v, w = self.data[index]
image_path = self.image_paths[image_name]
image = Image.open(image_path)
image = np.array(image)
image = image / 255
image = torch.tensor(image, dtype=torch.float32)
image = image.permute(2,0,1)
target=torch.tensor([v, w], dtype=torch.float32)
return image, target
FollowLineCNN.py
This file contains the FollowLineCNN class, which defines the custom CNN I used to train the model. I decided to use only three convolutional layers, changing a few parameters along the way for testing.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
import torch
import torch.nn as nn
import torch.optim as optim
class FollowLineCNN(nn.Module):
def __init__(self):
super().__init__()
self.conv1=nn.Conv2d(3, 16, kernel_size=3, stride=1, padding=1)
self.relu1=nn.ReLU()
self.pool1=nn.MaxPool2d(kernel_size=2, stride=2)
self.conv2=nn.Conv2d(16, 16, kernel_size=3, stride=1, padding=1)
self.relu2=nn.ReLU()
self.pool2=nn.MaxPool2d(kernel_size=2, stride=2)
self.conv3=nn.Conv2d(16, 16, kernel_size=3, stride=1, padding=1)
self.relu3=nn.ReLU()
self.pool3=nn.MaxPool2d(kernel_size=2, stride=2)
self.flatten=nn.Flatten()
self.linear=nn.Linear(76800, 2)
def forward(self, x):
x = self.conv1(x)
x = self.relu1(x)
x = self.pool1(x)
x = self.conv2(x)
x = self.relu2(x)
x = self.pool2(x)
x = self.conv3(x)
x = self.relu3(x)
x = self.pool3(x)
x = self.flatten(x)
x = self.linear(x)
return x
model_training_ai_follow_line.py
This is the main script for training and exporting the model. To balance the dataset, I used a weight-based system: each sample is weighted by the inverse frequency of its class, so images with an angular velocity of zero (the most common) get less weight and images with non-zero angular velocity get more. I exported the model to .onnx so it can be used inside the Unibotics platform, as the exercise statement indicates.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
import os
import matplotlib.pyplot as plt
import numpy as np
from PIL import Image
import torch
import torch.nn as nn
import torch.optim as optim
from FollowLineCNN import FollowLineCNN
from FollowLineDataset import FollowLineDataset
def datainfo(n_train_images, n_validation_images, v_values, w_values):
print(f"Total number of images: {len(n_train_images) + len(n_validation_images)}")
print(f"Total number of training samples: {len(n_train_images)}")
print(f"Min v: {min(v_values)}")
print(f"Max v: {max(v_values)}")
print(f"Mean v: {sum(v_values) / len(v_values):.4f}")
print(f"Min w: {min(w_values)}")
print(f"Max w: {max(w_values)}")
print(f"Mean w: {sum(w_values) / len(w_values):.4f}\n")
def hist_v_w(v_values, w_values):
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
plt.hist(w_values, bins=10, color="skyblue", density=True)
plt.xlabel("Value")
plt.ylabel("Frequency")
plt.title("W Histogram")
plt.subplot(1, 2, 2)
plt.hist(v_values, bins=10, color="skyblue", density=True)
plt.xlabel("Value")
plt.ylabel("Frequency")
plt.title("V Histogram")
plt.tight_layout()
plt.show()
def diagram_v_w(v_values, w_values):
plt.scatter(v_values, w_values)
plt.colorbar()
plt.show()
def dataimg():
dataset_path = "/media/hectormc/USB/Datasets_Jderobot/follow_line_dataset"
for actual_folder, _, files in os.walk(dataset_path):
if actual_folder == os.path.join(dataset_path, "test_images"):
continue
for file in files:
if not file.endswith(".png"):
continue
rute = os.path.join(actual_folder, file)
img = Image.open(rute)
print(img.size, img.mode)
def create_sample_weights(dataset):
w_values = np.array([sample[2] for sample in dataset.data], dtype=np.float32)
classes = np.digitize(w_values, bins=[-1.0, -0.2, 0.2, 1.0])
class_counts = np.bincount(classes, minlength=5)
print("\nClass distribution:")
for i, count in enumerate(class_counts):
print(f"Class {i}: {count} samples")
class_weights = 1.0 / class_counts
print("\nClass weights:")
for i, weight in enumerate(class_weights):
print(f"Class {i}: {weight:.8f}")
sample_weights = class_weights[classes]
return torch.tensor(sample_weights, dtype=torch.double)
def main():
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Device used: {device}")
if torch.cuda.is_available():
print(f"GPU: {torch.cuda.get_device_name(0)}")
base_path = "/media/hectormc/USB/Datasets_Jderobot/follow_line_dataset"
train_dataset = FollowLineDataset(os.path.join(base_path, "train.csv"))
validation_dataset = FollowLineDataset(os.path.join(base_path, "validation.csv"))
sample_weights = create_sample_weights(train_dataset)
sampler = torch.utils.data.WeightedRandomSampler(
weights=sample_weights,
num_samples=len(train_dataset),
replacement=True,
)
train_loader = torch.utils.data.DataLoader(
train_dataset, batch_size=16, sampler=sampler, num_workers=4
)
validation_loader = torch.utils.data.DataLoader(
validation_dataset, batch_size=16, shuffle=False, num_workers=4
)
model = FollowLineCNN().to(device)
loss_function = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)
n_epoch = 15
for epoch in range(n_epoch):
model.train()
train_loss = 0.0
for images, targets in train_loader:
images = images.to(device)
targets = targets.to(device)
predictions = model(images)
loss = loss_function(predictions, targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()
train_loss += loss.item()
train_loss /= len(train_loader)
model.eval()
validation_loss = 0.0
with torch.no_grad():
for images, targets in validation_loader:
images = images.to(device)
targets = targets.to(device)
predictions = model(images)
loss = loss_function(predictions, targets)
validation_loss += loss.item()
validation_loss /= len(validation_loader)
print(
f"Epoch: {epoch + 1}/{n_epoch} | "
f"Train Loss: {train_loss:.4f} | "
f"Validation Loss: {validation_loss:.4f}"
)
# Export to ONNX
model.eval()
model_cpu = model.to("cpu")
dummy_input = torch.randn(1, 3, 480, 640)
torch.onnx.export(
model_cpu,
dummy_input,
"follow_line.onnx",
input_names=["image"],
output_names=["output"],
)
print("\nLast epoch model exported:\nfollow_line.onnx")
if __name__ == "__main__":
main()
The last step to complete this task is testing the trained model in the exercise in Unibotics.
Visual Object Detection. Downloading the pre-trained model, the person dataset, and fine-tuning
While working on End to End Visual Control, I was also working on the Visual Object Detection exercise. Its objective is to detect people in front of the webcam and draw a bounding box around each one. To achieve this, the exercise requires fine-tuning a pre-existing model by following the guide for fine-tuning pre-existing models in PyTorch.
Over the last week I followed that fine-tuning tutorial. I downloaded the pre-trained model mobilenet-v1-ssd and part of the person dataset (only a subset of the images, because my device did not have enough space for the whole dataset). I am now fine-tuning the MobileNet model with the downloaded subset of the person dataset.