Migrate syscalls documentation from the website to Haiku internals book
Change-Id: I997049a1e174e7dd346e1342a2513a67e118a93b
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
The rc resource compiler
|
||||
====================
|
||||
========================
|
||||
|
||||
.. toctree::
|
||||
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 29 KiB |
@@ -1,5 +1,5 @@
|
||||
Device Driver Architecture
|
||||
==================================================
|
||||
==========================
|
||||
|
||||
This document tries to give you a short introduction into the new device
|
||||
manager, and how to write drivers for it. Haiku still supports the
|
||||
@@ -8,8 +8,8 @@ legacy device driver architecture introduced with BeOS.
|
||||
The new device driver architecture of Haiku is still a moving target,
|
||||
although most of its details are already specificed.
|
||||
|
||||
1. The Basics
|
||||
-------------
|
||||
The Basics
|
||||
----------
|
||||
|
||||
The device manager functionality builds upon *device_node* objects.
|
||||
Every driver in the system publishes one or more of such nodes, building
|
||||
@@ -35,8 +35,8 @@ bus.
|
||||
Every driver can also publish a device in */dev* for communication with
|
||||
userland applications. All drivers and devices are kernel modules.
|
||||
|
||||
2. Exploring the Device Tree
|
||||
----------------------------
|
||||
Exploring the Device Tree
|
||||
-------------------------
|
||||
|
||||
So how does it all work? When building the initial device tree, the
|
||||
system only explores a minimum of device drivers only, resulting in a
|
||||
@@ -64,8 +64,8 @@ directories "disk", "ports", and "bus", which will also allow to search
|
||||
matching drivers in "busses". While this is relatively limited, it is a
|
||||
good way to cut down the number of drivers to be loaded.
|
||||
|
||||
3. Writing a Driver
|
||||
-------------------
|
||||
Writing a Driver
|
||||
----------------
|
||||
|
||||
The device manager assumes the following API from a driver module:
|
||||
|
||||
@@ -128,8 +128,8 @@ either have dynamic children or fixed children, never both. Also, fixed
|
||||
children are registered before **register_child_devices()** is called,
|
||||
while dynamic children are registered afterwards.
|
||||
|
||||
4. Publishing a Device
|
||||
----------------------
|
||||
Publishing a Device
|
||||
-------------------
|
||||
|
||||
To publish a device entry in the device file system under */dev*, all
|
||||
your driver has to do is to call the
|
||||
@@ -176,8 +176,8 @@ A device module must export the following API:
|
||||
- **select()**
|
||||
- **deselect()**
|
||||
|
||||
5. Writing a Bus Driver
|
||||
-----------------------
|
||||
Writing a Bus Driver
|
||||
--------------------
|
||||
|
||||
A bus driver is a driver that represents a bus where one or more
|
||||
arbitrary devices can be attached to.
|
||||
@@ -192,8 +192,8 @@ exploration makes use of additional information the bus can provide in
|
||||
order to find a responsible device driver faster, and with less
|
||||
overhead.
|
||||
|
||||
5.1. Writing an Intelligent Bus Driver
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
Writing an Intelligent Bus Driver
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
If your bus knows what type of device is attached to, and also has
|
||||
vendor and device ID information about that device, it is considered to
|
||||
@@ -253,21 +253,21 @@ drivers that do not relate to a physical device: /dev/null, /dev/zero,
|
||||
/dev/random, etc. A "generic" bus has been added, and these drivers can
|
||||
attach to it.
|
||||
|
||||
6. Open Issues
|
||||
--------------
|
||||
Open Issues
|
||||
-----------
|
||||
|
||||
While most of the new device manager is fledged out, there are some
|
||||
areas that could use improvements or are problematic under certain
|
||||
requirements. Also, some parts just haven't been written yet.
|
||||
|
||||
6.1. generic/simple busses
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
generic/simple busses
|
||||
^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
6.2. Unpublishing
|
||||
^^^^^^^^^^^^^^^^^
|
||||
Unpublishing
|
||||
^^^^^^^^^^^^
|
||||
|
||||
6.4. Versioning
|
||||
^^^^^^^^^^^^^^^
|
||||
Versioning
|
||||
^^^^^^^^^^
|
||||
|
||||
The way the device manager works, it makes versioning of modules (which
|
||||
are supposed to be one of the strong points of the module system) much
|
||||
@@ -283,17 +283,17 @@ supports. We would then need a **request_version()** or
|
||||
that allows to specify the version of the parent node this device node
|
||||
wants to talk to.
|
||||
|
||||
6.5. Unregistering Nodes
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
Unregistering Nodes
|
||||
^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
6.6. Support for generic drivers is missing
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
Support for generic drivers is missing
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
This should probably be done by simply adding a simple bus driver named
|
||||
"generic" that generic drivers need to ask for.
|
||||
|
||||
6.7. Mappings, And Other Optimizations
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
Mappings, And Other Optimizations
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
Due to the way the device tree is built, the device manager could
|
||||
remember which driver served a given device node. That way, it wouldn't
|
||||
|
||||
@@ -4,6 +4,7 @@ Kernel
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
|
||||
/kernel/system_calls
|
||||
/kernel/device_manager_introduction
|
||||
/kernel/obsolete_pnp_manager
|
||||
/kernel/vm/swap_file_support
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 40 KiB |
@@ -0,0 +1,240 @@
|
||||
System calls
|
||||
============
|
||||
|
||||
This paper presents the implementation of system calls in haiku, and especially on x86 systems.
|
||||
The system call mechanism is what allows user land code to communicate with kernel land code.
|
||||
|
||||
The whole paper is focused on the example of one system call: is_computer_on. This api tells if the
|
||||
computer is currently powered on or not. Using this system call as a study is interesting because
|
||||
its implementation is quite simple, and it is a historical one in BeOS system (with its brother
|
||||
is_computer_on_fire, but it is not a system call in haiku :). However all elements presented here
|
||||
remain valid for any other system call.
|
||||
|
||||
Execution flow
|
||||
--------------
|
||||
|
||||
Let's start with an otherview of the global process invloved. Here is the whole excecution flow:
|
||||
|
||||
.. image:: syscall_bt.png
|
||||
:alt: syscall backtrace
|
||||
|
||||
When some code calls "is_computer_on", it will call this function from libroot. For this particular
|
||||
case, a direct call to the internal syscall api is done, that is calling _kern_is_computer_on. So
|
||||
in this case the only aim of "is_computer_on" is to be a public api to the system call. On other
|
||||
cases more things may be done before the "_kern_" call. All syscall functions name in user space
|
||||
begin with the "_kern_" prefix. "_kern_is_computer_on" is just a jump to a fixed location in the
|
||||
process address space called "commpage".
|
||||
|
||||
There can be different code living in the commpage. Depending on the cpu capabilities, either
|
||||
_user_syscall_sysenter or _user_syscall_int code is present at the same address. These two
|
||||
functions contain the few instructions needed to transfer execution context to the kernel.
|
||||
|
||||
After that the kernel code is executed. Depending on the cpu capabilities, either "x86_sysenter" or
|
||||
"trap99" will be called. Finally "handle_syscall" calls the handler corresponding to the syscall
|
||||
number via a callback located in the syscall table. In our example it is "_user_is_computer_on"
|
||||
whose huge code is:
|
||||
|
||||
.. code-block:: cpp
|
||||
|
||||
static inline int
|
||||
_user_is_computer_on(void)
|
||||
{
|
||||
return 1;
|
||||
}
|
||||
|
||||
|
||||
No let's see more in details each of these steps, by following the execution path.
|
||||
|
||||
Userland
|
||||
--------
|
||||
|
||||
libroot.so
|
||||
..........
|
||||
|
||||
Starting in userland, libroot is where everything is done: All internal ("_kern_xxx") APIs are
|
||||
located in it. However there is no source code for these functions. They are generated at
|
||||
compilation time by `gensyscalls <https://cgit.haiku-os.org/haiku/tree/src/tools/gensyscalls>`_ that
|
||||
is built as part of the Haiku compilation process. Gensyscalls uses the `private syscalls.h header file <https://cgit.haiku-os.org/haiku/tree/headers/private/system/syscalls.h>`_
|
||||
to generate a source assembly file containing all "_kern_xxx" definitions.
|
||||
The order of the function definitions in "syscalls.h" determine their syscall number. The generated
|
||||
assembly source is located at "generated/objects/haiku/x86/common/system/libroot/os/syscalls.S.inc".
|
||||
Here is what it looks like:
|
||||
|
||||
.. code-block:: cpp
|
||||
|
||||
SYSCALL0(_kern_is_computer_on, 0)
|
||||
SYSCALL4(_kern_generic_syscall, 1)
|
||||
SYSCALL2(_kern_getrlimit, 2)
|
||||
...
|
||||
|
||||
SYSCALLX are macros. The number after SYSCALL tells how many arguments are used in the system call.
|
||||
The two parameters of the macros are the function name, and the syscall number. The SYSCALLX macros
|
||||
are defined in `syscalls.inc <https://cgit.haiku-os.org/haiku/tree/src/system/libroot/os/arch/x86/syscalls.inc>`_. They contain code that sets the syscall number in eax,
|
||||
and jump to a fix location in the commpage.
|
||||
|
||||
The commpage
|
||||
............
|
||||
|
||||
The commpage is some pages of code that are mapped by the kernel in each loaded process at a fixed
|
||||
memory location. On the ia32 architecture, this address is 0xffff0000. The first part of this area
|
||||
is a pointer table of each function it contains. The second part of it is the actual code of the
|
||||
functions. This allows the kernel to let a process run optimized/specific code for the cpu in
|
||||
userland.
|
||||
|
||||
The syscall entry is the first of the commpage, so its offset is at 0xC from the start of the
|
||||
commpage. The content of the syscall function is filled in `x86_initialize_syscall in syscalls.cpp <https://cgit.haiku-os.org/haiku/tree/src/system/kernel/arch/x86/32/syscalls.cpp>`_. This is where the cpu capabilites are checked (for the syscall
|
||||
feature). Depending on them, either trap or sysenter mechanism is used. So depending on this, the
|
||||
code of "_user_syscall_int" or "_user_syscall_sysenter" is copied in the commpage. The definition
|
||||
of these two functions is in `syscalls_asm.S <https://cgit.haiku-os.org/haiku/tree/src/system/kernel/arch/x86/32/syscalls_asm.S>`_. The next figure shows the mapping of the commpage area:
|
||||
|
||||
.. image:: commpage.png
|
||||
:alt: commpage mapping
|
||||
|
||||
In the kernel
|
||||
-------------
|
||||
|
||||
system call handler
|
||||
...................
|
||||
|
||||
Just as there can be 2 codes in commpage to use a syscall, 2 handlers exist to trap them.
|
||||
"x86_sysenter" and "trap99" are both defined in `interrupts.S <https://cgit.haiku-os.org/haiku/tree/src/system/kernel/arch/x86/32/interrupts.S>`_. However both
|
||||
functions are always present in the kernel. This means that on a cpu that supports sysenter,
|
||||
calling syscalls with int99 is still working. Both functions set the cpu/system in the correct
|
||||
state before calling "handle_syscall".
|
||||
|
||||
handle_syscall
|
||||
..............
|
||||
|
||||
This function is the last one before the specific code of each syscall. It is also defined in
|
||||
"arch_interrupts.S". What it does is copy the syscall parameters to the correct place, and call the
|
||||
final handler. All these information are retrieved from the "syscall_info" structure defined in
|
||||
`ksyscalls.h <https://cgit.haiku-os.org/haiku/tree/headers/private/kernel/ksyscalls.h>`_
|
||||
|
||||
.. code-block:: cpp
|
||||
|
||||
typedef struct syscall_info {
|
||||
void *function; // pointer to the syscall function
|
||||
int parameter_size; // summed up parameter size
|
||||
} syscall_info;
|
||||
|
||||
The array object "kSyscallInfos" contains all definitions for all system calls of the system. The
|
||||
index in the array is the number of the system call. This object is also generated by gensyscall
|
||||
in "objects/haiku/x86/common/system/kernel/syscall_table.h". It typically looks like this:
|
||||
|
||||
.. code-block:: cpp
|
||||
|
||||
#define SYSCALL_COUNT 247
|
||||
|
||||
#ifndef _ASSEMBLER
|
||||
|
||||
const int kSyscallCount = SYSCALL_COUNT
|
||||
|
||||
const syscal_info kSyscallInfos[] = {
|
||||
{ (void *)_user_is_computer_on, 0},
|
||||
{ (void *)_user_generic_syscall, 16},
|
||||
{ (void *)_user_getrlimit, 8},
|
||||
...
|
||||
|
||||
|
||||
The first define is used by "handle_syscall" to check that the provided syscall number is correct.
|
||||
|
||||
Adding syscalls
|
||||
---------------
|
||||
|
||||
So now that all parts are explained, let's see the process of adding a new syscall.
|
||||
The steps to add a syscall are:
|
||||
|
||||
|
||||
- Add the syscall prototype to `syscalls.h <https://cgit.haiku-os.org/haiku/tree/headers/private/system/syscalls.h>`_. Name it _kern_(). Avoid adding unnecessary
|
||||
dependencies to the header. I.e. if your syscall has pointers to structs as arguments, there's no
|
||||
need to include the headers that define the structs.
|
||||
- Add a function prototype _user_() with the same signature as your syscall to a fitting kernel
|
||||
header under "headers/private/kernel/".
|
||||
- Make sure the header with the _user_() prototype is included by `syscalls.cpp <https://cgit.haiku-os.org/haiku/tree/src/system/kernel/syscalls.cpp>`_.
|
||||
- Implement _user_() in a fitting source file.
|
||||
|
||||
That's it. There are some general rules for the implementation of a syscall:
|
||||
|
||||
- If your syscall has a 64 bit return value (as opposed to the common 32 bit status_t/ssize_t/int
|
||||
etc.), call syscall_64_bit_return_value() at the very beginning.
|
||||
- Never access user memory directly or IOW, if your syscall has a parameter that is a pointer to
|
||||
something, never dereference the pointer. Also don't dereference pointers in structures you get
|
||||
from userland. If you have to access user data, first check that the pointer actually points to
|
||||
user address space, using the IS_USER_ADDRESS() macro (if not, fail with B_BAD_ADDRESS). Allocate
|
||||
kernel memory large enough to hold the user data. If it's a small structure or short string, use
|
||||
the stack, otherwise allocate on the heap. For variable sized data enforce maximum limits. Then
|
||||
copy the user data to your kernel memory using user_memcpy(). Parameters are the same as for
|
||||
memcpy(), but the return value is a status_t. If it's not B_OK, fail with B_BAD_ADDRESS. If you
|
||||
want to return a data structure to userland, use the same strategy (just with swapped parameters
|
||||
for user_memcpy(), of course).
|
||||
- If your syscall can block and can be interrupted, make it restartable (there are exceptions when
|
||||
that is not necessary/desired, but usually it is). Restartable means that if your syscall has be
|
||||
interrupted by a signal, the kernel can just invoke it again after the signal has been handled.
|
||||
It will get the exact same parameters, which in some cases requires some special handling. E.g.
|
||||
relative timeouts have to be converted to absolute ones and stored. There are inline functions in
|
||||
`syscall_restart.h <https://cgit.haiku-os.org/haiku/tree/headers/private/kernel/syscall_restart.h>` which help with that. If you don't have any problematic parameters,
|
||||
just invoke syscall_restart_handle_post() with B_INTERRUPTED before you return from the syscall,
|
||||
if the syscall has been interrupted. Most syscalls return error codes and the function returns
|
||||
the error code passed to it, so one can use it like "return syscall_restart_handle_post(error);".
|
||||
If you have to deal with relative timeouts, use the appropriate syscall_restart_handle_timeout_pre()
|
||||
function at the beginning and syscall_restart_handle_timeout_post() (instead of
|
||||
restart_handle_post()) at the end of the syscall. The latter stores the timeout for restart, the
|
||||
former converts the timeout to absolute, respectively restores the stored timeout on syscall
|
||||
restart.
|
||||
|
||||
|
||||
Hello world
|
||||
-----------
|
||||
|
||||
Let's finish this article with something I cannot resist to add here: An assembler hello world
|
||||
example using the write syscall. This is something that was already published a long time ago on
|
||||
<a href="http://asm.sourceforge.net/intro/hello.html#AEN159" rel="nofollow">linux assembly</a>.
|
||||
However this was for BeOS, and the syscall interface of haiku is not the same. We will use "int 99"
|
||||
to do the syscalls so that it works on all x86 systems. Looking in the generated "syscalls.S.inc"
|
||||
we can see that "write" is syscall number 131 and "exit" is syscall number 33. So here is the code:
|
||||
|
||||
.. code-block:: asm
|
||||
|
||||
section .text
|
||||
global _start
|
||||
|
||||
_syscall: ; system call
|
||||
int 99
|
||||
ret
|
||||
|
||||
_start: ; entry point for the linker
|
||||
push dword len ; message length
|
||||
push dword msg ; message to print
|
||||
push dword 0 ; 0 = offset (64 bits)
|
||||
push dword 0 ; 0 = offset (64 bits)
|
||||
push dword 1 ; 1 = stdout
|
||||
mov eax, 131 ; write syscall
|
||||
call _syscall
|
||||
add esp,20 ; restore stack
|
||||
|
||||
push dword 0 ; exit code
|
||||
mov eax, 33 ; exit syscall
|
||||
call _syscall
|
||||
|
||||
section .data
|
||||
|
||||
msg db "Hello world from Haiku syscall!",0xa
|
||||
len equ $ - msg
|
||||
|
||||
Yasm can be used to compile it, and ld to link it:
|
||||
|
||||
.. code-block:: sh
|
||||
|
||||
yasm -f elf hello.asm
|
||||
ld -s -o hello hello.o
|
||||
./hello
|
||||
Hello world from Haiku syscall!
|
||||
|
||||
|
||||
Conclusion
|
||||
----------
|
||||
|
||||
This was just a brief description of how system calls are processed. However it should help to see
|
||||
the global picture of this part of the Haiku system. There would be a lot of other things to tell
|
||||
around this subject, especially on the ia32 specific code, but this would go beyond the aim of this
|
||||
introduction paper.
|
||||
Reference in New Issue
Block a user