Serial read / write Best Practices

Hi there,

I'm trying to get a better grasp on best practices for data streams when sending data over serial. 2 major questions I am trying to answer are:

  • how can I best set a start / end bit/byte/bytes reliably
  • how can I ensure I'm not missing / dropping bytes or packets.

I've been doing a lot of digging and investigating and it looks like the most efficient way to send things is binary if possible. But it comes with the added challenge of knowing when the real data starts / stops and also how to decode it on the receivers end.

Let's say I'm trying to send pictures in raw format and want to send it in binary. How can I ensure my start / end bytes don't show up accidentally somewhere in the image data? Can I set serial to 7 bits and only use the 8th bit as marker for start / stop? Or is there a way to force the predefined bytes from showing up? ie, make FF into FE so FF can only mean start / stop? Or can I send like 4 sets of FF / 00? Or some other kind of predefined setup (I'm guessing this was solved decades ago but can't figure out what it is lol)

I believe the first 2 will work, but know the 2nd would... Except that means I can't use FF in data. The first seems like it might have even more limitations.

I'm also trying to figure out best practices for checksums but not sure where to start.

I saw one guy was sending 0-64 by converting them to the 0-o ASCII codes and converting that back to binary on the other end... I'm just not sure where to lookup how serial encodes / sends stats of bit streams, headers, etc....

Like I know there is an 8bit code for ASCII data, but how does the receiver know it's ASCII? Is there a header byte or bytes that serial sends to start the stream? So like if I send q as a string object, does serial send 8 bits (1 byte) only? Or does it send a header byte, the byte representing q and an end bytes?

If so do int or hex objects have headers and footers?

I know if I send a string of hex data is is received as b'/x##' where ## is the hex representation... which allows it to be decided.... But that means in order to send one byte of binary data, I need to send 7 bytes in string ASCII data for the hex representation of it.

Hint the hardware does not know the difference between Hex, Binary, Decimal or any other format, it is all 1's and 0's. The different types of data is for human understanding as we do not speak in ones or zeros but in sounds or there interpretation. We speak many languages and say the same thing in different ways but since the receiving end knows the language it is understood. The same is necessary for computers the receiver needs to be programed to receive and interpret per your code what it is receiving.

If you are sending a lot of code you can send it in packets. Your starting transmission tells how many bytes are coming, the last byte is the checksum, CRC or whatever you use. Fixed packets, the block of data will always be the same length, you just fill the unused area with dummy information. The receiver then notifies the sender it got the data and it checked or failed. If it is not acknowledged you get to try again or whatever you want to do. For short messages I prefer CAN (Controlled Area Network), it does all the testing etc for me saving lots of code.

The best practice is a subjective decision that is yours alone.

Thanks gilahultz,

Just checking... Are you saying that when serial is sending information it's not sending any headers, etc to indicate the encoding but just sending raw bytes and it's up to the other side to know how to decode it?

If that's the case, how does serial know the difference if only 1 byte is sent? It seems that serial is able to tell the difference between an int 5 and an ASCII 5... but is that just because I am putting the data into an int variable vs a string variable in Arduino? Like how does it know the difference between d and 100? (They both are the same binary 01100100)

The serial input basics tutorial may answer some of your questions.

As an example, you can send these data bytes using UART Port: 0x12, 0x34, 0x56, 0x78 in one of the following two ways:

1. Binary Code Transmission
"3-byte synchronization pattern: AC2B57 (for example)" "number of data bytes: here, it is: 4" "actual data bytes" "1-byte Check Sum computed from data bytes" (It is assumed that no consecutive 3-byte in the data stream will appear as AC2B57)

2. ASCII Code Transmission
"1-byte Message Start Mark: a control byte from the ASCII Table: STX (0x12)" "ASCII code for every digit of the data bytes: 0x31, 0x32, 0x33, 0x34, 0x35, 0x36, 0x37, 0x38" "1-byte Message End Mark: a control byte (Newline character or ETX or CR)"

That is the point it does not. In this case it is up to your software to determine what the bytes/bits etc are. Determining that is considered the software protocol you either chose or write. The hardware protocol controls the physical layer, not what is sent over it. The physical layer can change the data to what it can properly transmit and receive. There is a process where there is to many ones or zeros a bit of the opposite polarity is inserted, this is called bit stuffing. The receiving end understands what is going on and will pass on the correct data by removing the extra bit etc. CAN requires any receiver to acknowledge the transmission if it has been received up to a particular point by inserting a dominate bit in the packet.

Bit values determined by the receiver are determined by using a clock. With synchronous the clock information is sent with the data. How this information is sent is dependent on the medium used. Original the hardware was TTL and a clock of 16X the bit rate was used. This was easy do divide down in hardware to determine where to sample the bit. To give some tolerance sometimes extra bit time was inserted at the stop bit time, typically 1.5 or 2 bit times. Extra stops are treated as idle time and can be as long as required. The stop bit is a time thing, the bus is off for the stop time, same as idle time. Remains of this can be seen in the U(S)ART (Universal (Synchronous) Asynchronous Receiver Transmitter) of today.

In asynchronous the clock rate is fixed and already predefined in the receiver, for example: 9600 baud 9600 baud" means that the serial port is capable of transferring a maximum of 9600 bits per second. 115200 8N1 is the same but faster. If the information unit is one baud (one bit), then the bit rate and the baud rate are identical. In 'RS232' it is usually 8 bits of data, one stop no parity bit and one start bit (almost always a 1. You will see this is 9600 8N1. RS232 is a Recommended Standard #232 which defines the physical layer for communications for usage medium distances. Over the year it has evolved to what it is today, you can look up the standards as they have progressed over the years. Get the clocks wrong the data is garbage.

There is a lot of information on the web about this, and many tutorials etc. You could spend easily several days studying and looking this material up which I recommend you do. Hopefully this fills in many of the blanks.

Exactly. It is entirely up to you (ie, your software) to add any higher-level interpretations, protcols, etc.

It doesn't. It just sends the byte - it neither knows nor cares what that byte might mean to you.

What, exactly, do you mean by "serial" here?

Are you just referring to the hardware serial link, or are you thinking of the Arduino Serial library?
:thinking:

The Arduino Serial library does know about the C++ types of the data being passed to it - so it knows whether you are giving it an int to send or a string of char to send.

But, after the Library has done that processing, what it passes to the hardware is just a load of bytes - the hardware has no concept of where those bytes came from, nor what they might have meant to whatever sent them.

That example doesn't work - 'd' is not valid as a number!

If you wanted it to be interpreted as a number, you would have to write it as 0xd

The way that C++ recognises a number is:

  1. Context - ie, it is in a valid place for a number;
  2. It is an a valid format for a number.

https://en.cppreference.com/w/cpp/language/integer_literal

Probably Best Practice is not to reinvent the wheel - instead, use a well-established protocol.

You've posted in the PLC section, so I'm sure there must be well-established serial protocols in the PLC world ...

EDIT: I found a list of some protocols used in PLC world (probably not exhaustive; not all async serial):

Maybe this is what is confusing me....

It seems Arduino's Serial.print is only ever sending ASCII formatted bytes. Meaning if I send Serial.print(5,BIN) it is not sending a byte representing 5 in binary, it is sending 8 bytes representing the ASCII symbols for 00000101.

I've tried using Serial.write but it seems to clip any value I send over 255.... is there a built in function for converting and sending arrays of byte streams over serial? (meaning something that can handle more than 1 byte per index)

1st, I don't ever send information over PLC by serial and thus don't understand serial communications
2nd, we predefine a structure and send the whole array as bits if communicating between PLCs. All formatting is handled in the PLC
3rd, when we use someone's hardware (like a camera, etc). They provide an .eds / .esi file for ethernet / ethercat.

In none of these situations do I ever have to understand what is happening in headers, interpretation, etc. It just happens in the software.

ASCII d and the integer 100 are both written as 01100100, ie ASCII decoding interprets the integer 100 as the letter 'd'.

You can senb a byte array with
Serial.write(byteArray, numberOfBytes);

Thanks groundFungus,

I also figured out I can do this to split my ints into smaller bytes to send in a byte array... since an array of ints is clipping the first byte only.

void ShiftTo2Bytes(int i, byte *arr)
{
  arr[0] = (byte)(i & 0xFF);
  arr[1] = (byte)((i >> 8) & 0xFF);
}
int i = 1234;

Serial.write((byte*)&i, sizeof(i));

The (byte*) casts what follows to a pointer to (read: address of) a byte (array); the & gives you the address where i is located.

Indeed: the name "print" suggests that it is being sent to a "printer-like" device - ie, something that will display human-readable text.

OK - as I said, it was unclear whether you were using "serial" to mean the Arduino library (ie, C++ code), or the transmission hardware.

So, yes the character 'd' has the ASCII code value 100.

But, when written in C++ code (eg, an Arduino sketch), 'd' would not be valid as a number.

It all depends on context!

Yes - that is stated in the documentation:

If you want binary, use Serial.write instead:

But you asked what would be the best practice if you did - didn't you?

:man_shrugging:

@martiangnome
It is not only the print() method; there is something with the data type.

1.

byte y = 0x41;
Serial.print(y);   //what you expect to see on Serial Monitor?

2.

char y = 0x41;
Serial.print(y); //what you expect to see on Serial Monitor?