Tuesday, August 9, 2011

php error suppression

It is quite often that we can see error suppression operator @ is wildly used in legacy PHP application. For example:

@getimagesize($image);
@file($file);
@mysql_pconnect();
@mysql_query();
...

Let's see what PHP will do in the background when error suppression operator @ is used.

PHP will actually translate a simple statement @file($file) into three statements:

$saveOldSetting = error_reporting(0);
file($file);
error_reporting($saveOldSetting);

error_reporting(0) will trun off all error reporting and return the old error report setting value. Then PHP will execute file($file) as normal. After that, PHP will restore the error report setting to the old value.

We know what it does now. To use it or not, it is up to developer's judge. Personally, i believe every error should be handled decently so i won't use it in my code.

Monday, August 8, 2011

understanding PHP memory management

Let's keep exploring how PHP manages memory. Let's check this code and the output():

<?php
var_dump(memory_get_usage());
$name = 'henry';
var_dump(memory_get_usage());
unset($name);
var_dump(memory_get_usage());
?>

The output:

int(325496)
int(325672)
int(325532)

The first output tells us the memory usage is 325496. After we assign a value to a variable $name, the memory usage becomes 325672. The interesting part is the following code: unset($name),  which is supposed to release the memory taken by $name. But the output tells us, the memory usage after we unset($name) is 325532. But,
325496 - 325532 = -36. We still use 36 more memory even we unset($name)!

So, the question comes. Where does the 36 more memory go? Does unset truly release memory? Before we analyse this question in detail, I hope you stand firm with this answer: yes, unset really can release memory, no doubt about that(Surely it also depends on the zval.refcount, see this post: http://hengrui-li.blogspot.com/2011/08/php-copy-on-write-how-php-manages.html ).

Now let's see how PHP allocates memory. For a simple statement $name = 'henry', PHP will do at least two things: 1. PHP allocates memory for the name of the variable, which is '$name', and save it into a symbol table; 2. PHP allocates memory for the value of the variable, which is 'henry', and create a zval to save the value.

However, we must know this: When PHP is trying to ask memory from OS, it won't simply ask memory for its current task only. It will actually ask a large amount of memory(more than what it needs for the moment) from OS, and then allocate a small part of this 'large amount of memory'  to handle its current task. The benefit of asking large amount of memory at the beginning is, if PHP needs to allocate more memory later, it doesn't have to ask from OS again, and this avoids frequent system calling between PHP and OS.

When we do unset($name), PHP will release the memory. But it doesn't mean PHP will return the memory to OS. Actually, PHP will make the memory as its own spare memory that it can allocate to others if necessary.

Let's check this code:

<?php
var_dump(memory_get_usage(true));
$name = 'henry';
var_dump(memory_get_usage(true));
unset($name);
var_dump(memory_get_usage(true));
?>

Note memory_get_usage(true) means we want to get the real size of memory allocated from system. Now the output:

int(524288)
int(524288)
int(524288)

This output tells us one thing: when we do $name = 'henry', PHP does NOT ask more memory from OS.  

Now, let's answer the question at the beginning, where does the 36 more memory go? Let's check this code:

<?php
var_dump("dumping");
var_dump(memory_get_usage());
$name = 'henry';
var_dump(memory_get_usage());
unset($name);
var_dump(memory_get_usage());
?>

The output:

string(7) "dumping"
int(326668)
int(326808)
int(326668)

This time, 326668 - 326668 = 0. It is normal now. Memory allocation in PHP is quite implicitly, sometime it is hard for us to imagine. We var_dump("dumping"); at the beginning and everything looks normal now. This tells us, the 36 memory is taken by the output function. More precisely, the 36 memory is taken by the Header of the output.

Sunday, August 7, 2011

PHP copy on write - how PHP manages variable memory (2)

Continue with my last post http://hengrui-li.blogspot.com/2011/08/php-copy-on-write-how-php-manages.html, let's check some interesting cases.

case 1:

$name   = 'henry';
$fname = 'henry';

How does PHP handle these two variables?

//the output is: name: (refcount=1, is_ref=0)='henry'
xdebug_debug_zval('name');

// the output is:  fname: (refcount=1, is_ref=0)='henry'
xdebug_debug_zval('fname');

From the output, we can see that PHP actually creates two zval to store each of them.

case 2:

$fname = $name = 'henry';

// the output is: name: (refcount=2, is_ref=0)='henry'
xdebug_debug_zval('name');

//the output is: fname: (refcount=2, is_ref=0)='henry'
xdebug_debug_zval('fname');

In this case, we find that PHP only uses one zval, which is more efficient.

Thursday, August 4, 2011

PHP copy on write - how PHP manages variable memory

I've been asked a similar question a few times by a few developers so i think it is better to write it down. Let's check the code

//assume we have a large size array
$largeArray = getLargeSizeArray();

function doTask(Array $large)
{
                //do the task
}
doTask($largeArray);

The question is like this: the argument is passed by value, which means a copy of $largeArray is made. This will take a lot more memories. Is it better to pass the argument by reference: function doTask(Array &$large)?

To this question, my answer is always 'No'. Well, to be honest, i simply don't want developers think passing by reference is a good practice even for the sake of memory. But the truth is, it actually depends on what 'the task' is inside the doTask() function. Most of the time, we can simply pass by value.

To get a solid understanding, we better dig deeper. Let's see how Zend Engine manage variables internally.

Zend is actually using a C struct, zval, to store the value of a variable:

typedef struct _zval_struct {
    zvalue_value value;
    zend_uint refcount;
    zend_uchar type;
    zend_uchar is_ref;
  } zval;

'zvalue_value value' is where the value of the variable is stored. zvalue_value is a union:

typedef union _zvalue_value {
    long lval;
    double dval;
    struct {
        char *val;
        int len;
    } str;
    HashTable *ht;
    zend_object_value obj;
} zvalue_value;

As you can guess, zval.type stores the variable type. Zend is using this zval.type and zval.value to make PHP a weak typing language, even though C is strong type language. Anyway, typing is not what we want to discuss here.

A simple code: $name = 'henry'; The question is, how Zend uses a zval to store '$name', or, how a zval knows that it is storing a value for '$name'? In zval, we can't find any field to store the '$name'. The answer is, PHP stores the name of a variable in a hash table, called symbol_table. And there is a mapping mechanism from the variable name to the variable value(zval).

Now Let's check the PHP code:

$name  = 'henry';
$fname = $name;
unset($name);

The first line, PHP allocates a 6 bytes of memory to store 'henry'(5 bytes) and \0 (1 byte), which is NULL.
The second line, a new variable $fname is created, and the value of $name is "copied" to $fname.
The third line, unset $name trying to free the memory taken by $name.

This kind of code is quite common. If PHP allocates a new memory for every new variable assignment, then for this example, PHP must give 12 bytes of memory for $name and $fname. We know we don't really need that much memory. We can simply make symbol_table's $fname refers to the same zval that $name is referring to. And that is exactly how Zend Engine does. Humm, that sounds like we are not really copying, we are referring. So what will happen when we do unset($name)? How does PHP knows there is a $fname referring to $name?

Time to have a look at "zend_uint refcount". Let's try this:

$name='henry';
xdebug_debug_zval('name');

The output is "name: (refcount=1, is_ref=0)='henry'"

And we can see that zval.refcount=1, which means there is one variable referring to this zval. Now do this:

$fname = $name;
xdebug_debug_zval('fname');

The output is "fname: (refcount=2, is_ref=0)='henry'"! Strange? Shouldn't be refcount=1? Let's try:

xdebug_debug_zval('name');

We get the same output: "name: (refcount=2, is_ref=0)='henry'"!

So, actually, $fname and $name are referring to the same zval. By changing the value of zval.refcount, PHP knows that there are two variables referring to the same zval. If we assign $name to more other new variables, PHP will simply increase the value of zval.refcount and it will NOT allocate more memories. Ok, what happens if we unset($name)? You can guess! Right, PHP simply descrease the value of zval.refcount. Let's do this:

unset($name);
xdebug_debug_zval('fname');

The output is "fname: (refcount=1, is_ref=0)='henry'". You know what? In this situation(two variables referring to same zval), using unset cannot release/free the memory.

Now, try this:

$name  = 'henry';
$fname = $name;
$name  = 'li';

Obvious, $fname is still 'henry'. But if $fname is referring to the same zval, its value should change to 'li' too, right? Well, PHP has a copy on write mechanism: When PHP is going to change a variable, it will check its zval.refcount first. If zval.refcount > 1, PHP will create a new zval, descrease the old zval.refcount by 1, and modify the symbol_table so that $fname and $name is referring to different zval. So, at this time, PHP must allocate new memory. And also at this time, if we unset($name), we can really save some memory.

Now we know that when PHP is doing pass by value, or copying a variable to another, it is not "really copying". It makes them referring to the same zval to save memory. So back to the question at the beginning:

//assume we have a large size array
$largeArray = getLargeSizeArray();

function doTask(Array $large)
{
                //do the task
}
doTask($largeArray);

Simply passing the $largeArray by value into doTask function will not cost more memory. But, i also say it really depends on what we do inside the doTask() function. If we need to change the value of the argument, then PHP has to spend more memory. Like this:

//assume we have a large size array
$largeArray = getLargeSizeArray();

function doTask(Array $large)
{
                $large[0] = 'xxxx';
}
doTask($largeArray);

We change the value of the argument and PHP has to create a new zval to save it.

Alright, finally, just simply mention it here: what is "zend_uchar is_ref"? I think you can easily guess now:

$name  = 'henry';
$fname = &$name;
xdebug_debug_zval('name');

The output is "name: (refcount=2, is_ref=1)='henry'". Don't have to explain more, right?

Wednesday, August 3, 2011

lovely Python (compared with PHP or maybe some other languages)

The Zen of Python, especially "There should be one - and preferably only one - obvious way to do it", highly attracts me. And that actually drives me to start learning Python. And once I start, I find I like it more. Just some simple examples can show how Python practise what it preaches.

Python code

name = 'henry'
if name == 'hengrui':
            print ('Hello ' + name)
else:
            print ('Is ' + name + ' your real name?')

PHP code
version 1

$name = 'henry';
if  ($name == 'hengrui') {
            echo 'Hello ', $name;
} else {
            echo 'Is ', $name, ' your real name?';
}

version 2

$name = 'henry';
if  ($name === 'hengrui') {
            echo 'Hello ', $name;
} else {
            echo 'Is ', $name, ' your real name?';
}
Found the difference in this version? Well, it is '===' not '=='.

version 3

$name = 'henry'
if  ($name == 'hengrui')
            echo 'Hello ', $name;
else
            echo 'Is ', $name, ' your real name?';

version 4

$name = 'henry'
if  ($name == 'hengrui') echo 'Hello ', $name;
else
            echo 'Is ', $name, ' your real name?';

Surely we can still have some other versions of implementation in PHP code. Like it or not? It is up to you. But for me, i don't like this kind of variations and definitely coding conventions must be set up to ensure we won't have all these different coding practices in serious projects.

In my post, http://hengrui-li.blogspot.com/2011/05/variable-assignment-inside-if.html, i believe doing assignment in if statement is a very bad practice. In Python, you just can't do that.
PHP code

function getResult()
{
                //oh, by the way, for boolean value, you can also return TRUE, True. Anyway, case insensitive
                return true;
}
//this code works in PHP
if ($result = getResult()) {
                echo 'it is true';
}

Python code
def getResult():
                #you can only use True, case sensitive.
                return True
if result = getResult():
                print('it is true')

The python code simply cannot work. It doesn't allows you to do assignment in if statement. in a if statement, the only one thing you should do and you can do is logic operation.

Another very simple example:
PHP code

//both works
print 'hello';
print ('hello');

Python code

#this is correct
print ('hello')
#this is syntax error in python 3
print 'hello'

You may say, hey in this case PHP is better because you can type less. Well, typing more or less doesn't matter here. The point is: "There should be one - and preferably only one - obvious way to do it". (Actually, in python 2, print is a statement as well and it works exactly like PHP print. But Python 3 corrects this issue, which is good!)

Tuesday, August 2, 2011

javascript unit test with qunit

Qunit, http://docs.jquery.com/Qunit, is a javascript test suite. It is used by jQuery project for testing. But we can use it to test our javascript code as well.

Firstly, we need to setup our testing environment. Let's create a file called testSuite.html

<html>
<head>
<link rel="stylesheet" href="qunit.css" type="text/css" media="screen">
<script type="text/javascript" src="qunit.js"></script>
</head>
<body>
<h1 id="qunit-header">QUnit Test Suite</h1>
<h2 id="qunit-banner"></h2>
<div id="qunit-testrunner-toolbar"></div>
<h2 id="qunit-userAgent"></h2>
<ol id="qunit-tests"></ol>
</body>
</html>

Secondly, we must include our javascript code that needs to be tested.For example, if we put our javascript code in app.js file, we need to include this file in testSuite.html.

<html>
<head>
<link rel="stylesheet" href="qunit.css" type="text/css" media="screen">
<script type="text/javascript" src="qunit.js"></script>
<!-- The files that include the js code for testing -->
<script type="text/javascript" src="app.js"></script>
</head>
<body>
<h1 id="qunit-header">QUnit Test Suite</h1>
<h2 id="qunit-banner"></h2>
<div id="qunit-testrunner-toolbar"></div>
<h2 id="qunit-userAgent"></h2>
<ol id="qunit-tests"></ol>
</body>
</html>

Thirdly, we need to write test cases. Assuming we put test cases in testCases.js, we need to include this file in testSuite.html too.

<html>
<head>
<link rel="stylesheet" href="qunit.css" type="text/css" media="screen">
<script type="text/javascript" src="qunit.js"></script>
<!-- The files that include the js code for testing -->
<script type="text/javascript" src="app.js"></script>
<!-- The files that include the test cases -->
<script type="text/javascript" src="testCases.js"></script>
</head>
<body>
<h1 id="qunit-header">QUnit Test Suite</h1>
<h2 id="qunit-banner"></h2>
<div id="qunit-testrunner-toolbar"></div>
<h2 id="qunit-userAgent"></h2>
<ol id="qunit-tests"></ol>
</body>
</html>

Now, suppose we have this simple function in app.js:

function isLarger(a, b) {
                return a > b;
}

In testCases.js, we have our test case:

test('isLarger()', function() {
//success
ok(isLarger(3,1), "3 is larger than 1");
// fails
ok(isLarger(3, 5), '3 is larger than 5');
});

Ok, it is time to run our tests. try http://localhost/testSuite.html, and the result is like below:

Looks nice, doesn't it?

However, you may wonder, most of our javascript code is about DOM manipulation, how to test that? Let 's have a look.

DOM manipulation test
Suppose we have a function in app.js:
function domManipulation()
{
            var dom = document.getElementById('input');
            dom.value = 'henry';
}

Obviously, to run this function correctly, we need to setup our DOM document for testing. Fortunately, Qunit provide us with #qunit-fixture element. Let's change our testSuite.html a bit:

<html>
<head>
<link rel="stylesheet" href="qunit.css" type="text/css" media="screen">
<script type="text/javascript" src="qunit.js"></script>
<!-- The files that include the js code for testing -->
<script type="text/javascript" src="app.js"></script>
<!-- The files that include the test cases -->
<script type="text/javascript" src="testCases.js"></script>
</head>
<body>
<h1 id="qunit-header">QUnit Test Suite</h1>
<h2 id="qunit-banner"></h2>
<div id="qunit-testrunner-toolbar"></div>
<h2 id="qunit-userAgent"></h2>
<ol id="qunit-tests"></ol>
<!-- The DOM elements for testing -->
<div id="qunit-fixture">
<input id="input" type="text" value="change me" />
</div>
</body>
</html>

And then, add a test case in our testCases.js:
test('domManipulation', function(){
                domManipulation();
                var dom = document.getElementById('input');
                //success
                equals('henry', dom.value, 'input value changed!');
               //failure
                equals('change me', dom.value, 'input value not changed!');
});

Now, refresh our testSuite.html, and the result is like below:


PS: Qunit is a javascript test suite. We know that another popular frontend test suite is Selenium, which is so cool as well. It definitely worth your time on it.

Monday, August 1, 2011

php invoke method

First of all, I'm not suggesting using this magic method in serious projects. I simply want to investigate this interesting method and see what it can do.

PHP 5.3 introduced this '__invoke' magic method. From PHP manual, "The __invoke method is called when a script tries to call an object as a function". Let's check this code:

class FirstClassFunction
{
                public function __invoke()
                {
                                var_dump(func_get_args());
                                var_dump("i was called");
                }
}

$f = new FirstClassFunction();
$f(1,2,3);

$callback = $f;
$callback('one','two', 'three');

call_user_func($callback);

Amazing feature, isn't it? It enables PHP to accommodate pseudo-first-class functions, somehow. We can probably also use this feature for tasks like passing a function around (a little like anonymous function?).

In the world of PHP, it seems wierd that something can be an object and a function at the same time. But this couldn't be more familiar to a javascript developer. In javascript, function is first-class object. But let's go back to PHP's OO system. Usually, when we try to define or model a class, the first thing to do is define this class's role, its responsibilities and jobs in business domain. So usually a domain class represents something. But quite often, we also write some classes that don't represent anything. A Utility class is a typical example. A Utility class usually contains some static helper functions. It is totally fine that we don't wrap these functions into a class but just use them as normal procedure functions. Check this code:

//put an escape function into Utility class
$text = Utility::escape($text);

//or we can do it the other way
//we put all helper functions into helpers.php
require 'helpers.php';
$text = escape($text);

For this example, Utility is simply a namespace for helper functions. Also, if we want to utilise autoloading instead of 'require', we need to put helper functions into a class. I personally would prefer wrapping everything into classes. But it is fine if some developers just want to use these kind of functions in procedure way. If I'm not wrong, CodeIgniter impelment its helper functions as procedure functions, although i'm not familiar with CodeIgniter.

Ok, Utility class does not represent anything but simply provides a group of helper functions. We also have some classes that are suppose to do only single job. Sometimes, these classes are named with verbs like a Login class, or sometimes not verbs but obviouslly some kind of action executer like a Router class. A Router class should only do route and a Login class should only do login, nothing else. Usually, what we do is we create some default methods for these classes. Say for example, run(), exec(), whatever and the code is like:

$login = new Login();
$login->run();
$login->exec();
$login->login();

Now, with the __invoke method, we can do:
$login = new Login();
$login();

From the readability perspective, is it worse than using default methods? I think it depends on different developers. Some might think __invoke is a nice feature but others might think it is a huge headache.

But, as i stated in the beginning, i'm not suggesting using it in serious projects. Every magic method does bring some costs and confusions and __invoke is not excluded. One issue with __invoke is if your object is a result of a function call, you have to save it into a variable first. This won't work: ObjectMaker::factory()(). You have to do $obj = ObjectMaker::factory(); $obj(); Another issue is, in my opinion, it does somehow make the code less clean and confusing to PHP developers. People will think $obj() is a function and don't know that it is actually an object.