Files
docs-rust/ch03/ch03-02-data-types.html
2026-06-22 21:27:36 +05:30

383 lines
22 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Data Types</title>
</head>
<body>
<h2 id="data-types"><a class="header" href="#data-types">Data Types</a></h2>
<p>Every value in Rust is of a certain <em>data type</em>, which tells Rust what kind of
data is being specified so that it knows how to work with that data. Well look
at two data type subsets: scalar and compound.</p>
<p>Keep in mind that Rust is a <em>statically typed</em> language, which means that it
must know the types of all variables at compile time. The compiler can usually
infer what type we want to use based on the value and how we use it. In cases
when many types are possible, such as when we converted a <code>String</code> to a numeric
type using <code>parse</code> in the <a href="../ch02/ch02-00-guessing-game-tutorial.html#comparing-the-guess-to-the-secret-number">“Comparing the Guess to the Secret
Number”</a><!-- ignore --> section in
Chapter 2, we must add a type annotation, like this:</p>
<pre class="playground"><code class="language-rust edition2024"><span class="boring">#![allow(unused)]
</span><span class="boring">fn main() {
</span>let guess: u32 = "42".parse().expect("Not a number!");
<span class="boring">}</span></code></pre>
<p>If we dont add the <code>: u32</code> type annotation shown in the preceding code, Rust
will display the following error, which means the compiler needs more
information from us to know which type we want to use:</p>
<pre><code class="language-console">$ cargo build
Compiling no_type_annotations v0.1.0 (file:///projects/no_type_annotations)
error[E0284]: type annotations needed
--&gt; src/main.rs:2:9
|
2 | let guess = "42".parse().expect("Not a number!");
| ^^^^^ ----- type must be known at this point
|
= note: cannot satisfy `&lt;_ as FromStr&gt;::Err == _`
help: consider giving `guess` an explicit type
|
2 | let guess: /* Type */ = "42".parse().expect("Not a number!");
| ++++++++++++
For more information about this error, try `rustc --explain E0284`.
error: could not compile `no_type_annotations` (bin "no_type_annotations") due to 1 previous error
</code></pre>
<p>Youll see different type annotations for other data types.</p>
<h3 id="scalar-types"><a class="header" href="#scalar-types">Scalar Types</a></h3>
<p>A <em>scalar</em> type represents a single value. Rust has four primary scalar types:
integers, floating-point numbers, Booleans, and characters. You may recognize
these from other programming languages. Lets jump into how they work in Rust.</p>
<h4 id="integer-types"><a class="header" href="#integer-types">Integer Types</a></h4>
<p>An <em>integer</em> is a number without a fractional component. We used one integer
type in Chapter 2, the <code>u32</code> type. This type declaration indicates that the
value its associated with should be an unsigned integer (signed integer types
start with <code>i</code> instead of <code>u</code>) that takes up 32 bits of space. Table 3-1 shows
the built-in integer types in Rust. We can use any of these variants to declare
the type of an integer value.</p>
<p><span class="caption">Table 3-1: Integer Types in Rust</span></p>
<div class="table-wrapper">
<table>
<thead>
<tr><th>Length</th><th>Signed</th><th>Unsigned</th></tr>
</thead>
<tbody>
<tr><td>8-bit</td><td><code>i8</code></td><td><code>u8</code></td></tr>
<tr><td>16-bit</td><td><code>i16</code></td><td><code>u16</code></td></tr>
<tr><td>32-bit</td><td><code>i32</code></td><td><code>u32</code></td></tr>
<tr><td>64-bit</td><td><code>i64</code></td><td><code>u64</code></td></tr>
<tr><td>128-bit</td><td><code>i128</code></td><td><code>u128</code></td></tr>
<tr><td>Architecture-dependent</td><td><code>isize</code></td><td><code>usize</code></td></tr>
</tbody>
</table>
</div>
<p>Each variant can be either signed or unsigned and has an explicit size.
<em>Signed</em> and <em>unsigned</em> refer to whether its possible for the number to be
negative—in other words, whether the number needs to have a sign with it
(signed) or whether it will only ever be positive and can therefore be
represented without a sign (unsigned). Its like writing numbers on paper: When
the sign matters, a number is shown with a plus sign or a minus sign; however,
when its safe to assume the number is positive, its shown with no sign.
Signed numbers are stored using <a href="https://en.wikipedia.org/wiki/Two%27s_complement">twos complement</a><!-- ignore
--> representation.</p>
<p>Each signed variant can store numbers from (2<sup>n 1</sup>) to 2<sup>n
1</sup> 1 inclusive, where <em>n</em> is the number of bits that variant uses. So, an
<code>i8</code> can store numbers from (2<sup>7</sup>) to 2<sup>7</sup> 1, which equals
128 to 127. Unsigned variants can store numbers from 0 to 2<sup>n</sup> 1,
so a <code>u8</code> can store numbers from 0 to 2<sup>8</sup> 1, which equals 0 to 255.</p>
<p>Additionally, the <code>isize</code> and <code>usize</code> types depend on the architecture of the
computer your program is running on: 64 bits if youre on a 64-bit architecture
and 32 bits if youre on a 32-bit architecture.</p>
<p>You can write integer literals in any of the forms shown in Table 3-2. Note
that number literals that can be multiple numeric types allow a type suffix,
such as <code>57u8</code>, to designate the type. Number literals can also use <code>_</code> as a
visual separator to make the number easier to read, such as <code>1_000</code>, which will
have the same value as if you had specified <code>1000</code>.</p>
<p><span class="caption">Table 3-2: Integer Literals in Rust</span></p>
<div class="table-wrapper">
<table>
<thead>
<tr><th>Number literals</th><th>Example</th></tr>
</thead>
<tbody>
<tr><td>Decimal</td><td><code>98_222</code></td></tr>
<tr><td>Hex</td><td><code>0xff</code></td></tr>
<tr><td>Octal</td><td><code>0o77</code></td></tr>
<tr><td>Binary</td><td><code>0b1111_0000</code></td></tr>
<tr><td>Byte (<code>u8</code> only)</td><td><code>b'A'</code></td></tr>
</tbody>
</table>
</div>
<p>So how do you know which type of integer to use? If youre unsure, Rusts
defaults are generally good places to start: Integer types default to <code>i32</code>.
The primary situation in which youd use <code>isize</code> or <code>usize</code> is when indexing
some sort of collection.</p>
<section class="note" aria-role="note">
<h5 id="integer-overflow"><a class="header" href="#integer-overflow">Integer Overflow</a></h5>
<p>Lets say you have a variable of type <code>u8</code> that can hold values between 0 and
255. If you try to change the variable to a value outside that range, such as
256, <em>integer overflow</em> will occur, which can result in one of two behaviors.
When youre compiling in debug mode, Rust includes checks for integer overflow
that cause your program to <em>panic</em> at runtime if this behavior occurs. Rust
uses the term <em>panicking</em> when a program exits with an error; well discuss
panics in more depth in the <a href="../ch09/ch09-01-unrecoverable-errors-with-panic.html">“Unrecoverable Errors with
<code>panic!</code></a><!-- ignore --> section in Chapter
9.</p>
<p>When youre compiling in release mode with the <code>--release</code> flag, Rust does
<em>not</em> include checks for integer overflow that cause panics. Instead, if
overflow occurs, Rust performs <em>twos complement wrapping</em>. In short, values
greater than the maximum value the type can hold “wrap around” to the minimum
of the values the type can hold. In the case of a <code>u8</code>, the value 256 becomes
0, the value 257 becomes 1, and so on. The program wont panic, but the
variable will have a value that probably isnt what you were expecting it to
have. Relying on integer overflows wrapping behavior is considered an error.</p>
<p>To explicitly handle the possibility of overflow, you can use these families
of methods provided by the standard library for primitive numeric types:</p>
<ul>
<li>Wrap in all modes with the <code>wrapping_*</code> methods, such as <code>wrapping_add</code>.</li>
<li>Return the <code>None</code> value if there is overflow with the <code>checked_*</code> methods.</li>
<li>Return the value and a Boolean indicating whether there was overflow with
the <code>overflowing_*</code> methods.</li>
<li>Saturate at the values minimum or maximum values with the <code>saturating_*</code>
methods.</li>
</ul>
</section>
<h4 id="floating-point-types"><a class="header" href="#floating-point-types">Floating-Point Types</a></h4>
<p>Rust also has two primitive types for <em>floating-point numbers</em>, which are
numbers with decimal points. Rusts floating-point types are <code>f32</code> and <code>f64</code>,
which are 32 bits and 64 bits in size, respectively. The default type is <code>f64</code>
because on modern CPUs, its roughly the same speed as <code>f32</code> but is capable of
more precision. All floating-point types are signed.</p>
<p>Heres an example that shows floating-point numbers in action:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let x = 2.0; // f64
let y: f32 = 3.0; // f32
}</code></pre>
<p>Floating-point numbers are represented according to the IEEE-754 standard.</p>
<h4 id="numeric-operations"><a class="header" href="#numeric-operations">Numeric Operations</a></h4>
<p>Rust supports the basic mathematical operations youd expect for all the number
types: addition, subtraction, multiplication, division, and remainder. Integer
division truncates toward zero to the nearest integer. The following code shows
how youd use each numeric operation in a <code>let</code> statement:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
// addition
let sum = 5 + 10;
// subtraction
let difference = 95.5 - 4.3;
// multiplication
let product = 4 * 30;
// division
let quotient = 56.7 / 32.2;
let truncated = -5 / 3; // Results in -1
// remainder
let remainder = 43 % 5;
}</code></pre>
<p>Each expression in these statements uses a mathematical operator and evaluates
to a single value, which is then bound to a variable. <a href="../appendix/appendix-02-operators.html">Appendix
B</a><!-- ignore --> contains a list of all operators that Rust
provides.</p>
<h4 id="the-boolean-type"><a class="header" href="#the-boolean-type">The Boolean Type</a></h4>
<p>As in most other programming languages, a Boolean type in Rust has two possible
values: <code>true</code> and <code>false</code>. Booleans are one byte in size. The Boolean type in
Rust is specified using <code>bool</code>. For example:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let t = true;
let f: bool = false; // with explicit type annotation
}</code></pre>
<p>The main way to use Boolean values is through conditionals, such as an <code>if</code>
expression. Well cover how <code>if</code> expressions work in Rust in the <a href="ch03-05-control-flow.html#control-flow">“Control
Flow”</a><!-- ignore --> section.</p>
<h4 id="the-character-type"><a class="header" href="#the-character-type">The Character Type</a></h4>
<p>Rusts <code>char</code> type is the languages most primitive alphabetic type. Here are
some examples of declaring <code>char</code> values:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let c = 'z';
let z: char = ''; // with explicit type annotation
let heart_eyed_cat = '😻';
}</code></pre>
<p>Note that we specify <code>char</code> literals with single quotation marks, as opposed to
string literals, which use double quotation marks. Rusts <code>char</code> type is 4
bytes in size and represents a Unicode scalar value, which means it can
represent a lot more than just ASCII. Accented letters; Chinese, Japanese, and
Korean characters; emojis; and zero-width spaces are all valid <code>char</code> values in
Rust. Unicode scalar values range from <code>U+0000</code> to <code>U+D7FF</code> and <code>U+E000</code> to
<code>U+10FFFF</code> inclusive. However, a “character” isnt really a concept in Unicode,
so your human intuition for what a “character” is may not match up with what a
<code>char</code> is in Rust. Well discuss this topic in detail in <a href="../ch08/ch08-02-strings.html#storing-utf-8-encoded-text-with-strings">“Storing UTF-8
Encoded Text with Strings”</a><!-- ignore --> in Chapter 8.</p>
<h3 id="compound-types"><a class="header" href="#compound-types">Compound Types</a></h3>
<p><em>Compound types</em> can group multiple values into one type. Rust has two
primitive compound types: tuples and arrays.</p>
<h4 id="the-tuple-type"><a class="header" href="#the-tuple-type">The Tuple Type</a></h4>
<p>A <em>tuple</em> is a general way of grouping together a number of values with a
variety of types into one compound type. Tuples have a fixed length: Once
declared, they cannot grow or shrink in size.</p>
<p>We create a tuple by writing a comma-separated list of values inside
parentheses. Each position in the tuple has a type, and the types of the
different values in the tuple dont have to be the same. Weve added optional
type annotations in this example:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let tup: (i32, f64, u8) = (500, 6.4, 1);
}</code></pre>
<p>The variable <code>tup</code> binds to the entire tuple because a tuple is considered a
single compound element. To get the individual values out of a tuple, we can
use pattern matching to destructure a tuple value, like this:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let tup = (500, 6.4, 1);
let (x, y, z) = tup;
println!("The value of y is: {y}");
}</code></pre>
<p>This program first creates a tuple and binds it to the variable <code>tup</code>. It then
uses a pattern with <code>let</code> to take <code>tup</code> and turn it into three separate
variables, <code>x</code>, <code>y</code>, and <code>z</code>. This is called <em>destructuring</em> because it breaks
the single tuple into three parts. Finally, the program prints the value of
<code>y</code>, which is <code>6.4</code>.</p>
<p>We can also access a tuple element directly by using a period (<code>.</code>) followed by
the index of the value we want to access. For example:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let x: (i32, f64, u8) = (500, 6.4, 1);
let five_hundred = x.0;
let six_point_four = x.1;
let one = x.2;
}</code></pre>
<p>This program creates the tuple <code>x</code> and then accesses each element of the tuple
using their respective indices. As with most programming languages, the first
index in a tuple is 0.</p>
<p>The tuple without any values has a special name, <em>unit</em>. This value and its
corresponding type are both written <code>()</code> and represent an empty value or an
empty return type. Expressions implicitly return the unit value if they dont
return any other value.</p>
<h4 id="the-array-type"><a class="header" href="#the-array-type">The Array Type</a></h4>
<p>Another way to have a collection of multiple values is with an <em>array</em>. Unlike
a tuple, every element of an array must have the same type. Unlike arrays in
some other languages, arrays in Rust have a fixed length.</p>
<p>We write the values in an array as a comma-separated list inside square
brackets:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let a = [1, 2, 3, 4, 5];
}</code></pre>
<p>Arrays are useful when you want your data allocated on the stack, the same as
the other types we have seen so far, rather than the heap (we will discuss the
stack and the heap more in <a href="../ch04/ch04-01-what-is-ownership.html#the-stack-and-the-heap">Chapter 4</a><!-- ignore -->) or when
you want to ensure that you always have a fixed number of elements. An array
isnt as flexible as the vector type, though. A vector is a similar collection
type provided by the standard library that <em>is</em> allowed to grow or shrink in
size because its contents live on the heap. If youre unsure whether to use an
array or a vector, chances are you should use a vector. <a href="../ch08/ch08-01-vectors.html">Chapter
8</a><!-- ignore --> discusses vectors in more detail.</p>
<p>However, arrays are more useful when you know the number of elements will not
need to change. For example, if you were using the names of the month in a
program, you would probably use an array rather than a vector because you know
it will always contain 12 elements:</p>
<pre class="playground"><code class="language-rust edition2024"><span class="boring">#![allow(unused)]
</span><span class="boring">fn main() {
</span>let months = ["January", "February", "March", "April", "May", "June", "July",
"August", "September", "October", "November", "December"];
<span class="boring">}</span></code></pre>
<p>You write an arrays type using square brackets with the type of each element,
a semicolon, and then the number of elements in the array, like so:</p>
<pre class="playground"><code class="language-rust edition2024"><span class="boring">#![allow(unused)]
</span><span class="boring">fn main() {
</span>let a: [i32; 5] = [1, 2, 3, 4, 5];
<span class="boring">}</span></code></pre>
<p>Here, <code>i32</code> is the type of each element. After the semicolon, the number <code>5</code>
indicates the array contains five elements.</p>
<p>You can also initialize an array to contain the same value for each element by
specifying the initial value, followed by a semicolon, and then the length of
the array in square brackets, as shown here:</p>
<pre class="playground"><code class="language-rust edition2024"><span class="boring">#![allow(unused)]
</span><span class="boring">fn main() {
</span>let a = [3; 5];
<span class="boring">}</span></code></pre>
<p>The array named <code>a</code> will contain <code>5</code> elements that will all be set to the value
<code>3</code> initially. This is the same as writing <code>let a = [3, 3, 3, 3, 3];</code> but in a
more concise way.</p>
<!-- Old headings. Do not remove or links may break. -->
<p><a id="accessing-array-elements"></a></p>
<h4 id="array-element-access"><a class="header" href="#array-element-access">Array Element Access</a></h4>
<p>An array is a single chunk of memory of a known, fixed size that can be
allocated on the stack. You can access elements of an array using indexing,
like this:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre class="playground"><code class="language-rust edition2024">fn main() {
let a = [1, 2, 3, 4, 5];
let first = a[0];
let second = a[1];
}</code></pre>
<p>In this example, the variable named <code>first</code> will get the value <code>1</code> because that
is the value at index <code>[0]</code> in the array. The variable named <code>second</code> will get
the value <code>2</code> from index <code>[1]</code> in the array.</p>
<h4 id="invalid-array-element-access"><a class="header" href="#invalid-array-element-access">Invalid Array Element Access</a></h4>
<p>Lets see what happens if you try to access an element of an array that is past
the end of the array. Say you run this code, similar to the guessing game in
Chapter 2, to get an array index from the user:</p>
<p><span class="filename">Filename: src/main.rs</span></p>
<pre><code class="language-rust ignore panics">use std::io;
fn main() {
let a = [1, 2, 3, 4, 5];
println!("Please enter an array index.");
let mut index = String::new();
io::stdin()
.read_line(&amp;mut index)
.expect("Failed to read line");
let index: usize = index
.trim()
.parse()
.expect("Index entered was not a number");
let element = a[index];
println!("The value of the element at index {index} is: {element}");
}</code></pre>
<p>This code compiles successfully. If you run this code using <code>cargo run</code> and
enter <code>0</code>, <code>1</code>, <code>2</code>, <code>3</code>, or <code>4</code>, the program will print out the corresponding
value at that index in the array. If you instead enter a number past the end of
the array, such as <code>10</code>, youll see output like this:</p>
<!-- manual-regeneration
cd listings/ch03-common-programming-concepts/no-listing-15-invalid-array-access
cargo run
10
-->
<pre><code class="language-console">thread 'main' panicked at src/main.rs:19:19:
index out of bounds: the len is 5 but the index is 10
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
</code></pre>
<p>The program resulted in a runtime error at the point of using an invalid
value in the indexing operation. The program exited with an error message and
didnt execute the final <code>println!</code> statement. When you attempt to access an
element using indexing, Rust will check that the index youve specified is less
than the array length. If the index is greater than or equal to the length,
Rust will panic. This check has to happen at runtime, especially in this case,
because the compiler cant possibly know what value a user will enter when they
run the code later.</p>
<p>This is an example of Rusts memory safety principles in action. In many
low-level languages, this kind of check is not done, and when you provide an
incorrect index, invalid memory can be accessed. Rust protects you against this
kind of error by immediately exiting instead of allowing the memory access and
continuing. Chapter 9 discusses more of Rusts error handling and how you can
write readable, safe code that neither panics nor allows invalid memory access.</p>
</body>
</html>